Tiny LLM Studio — Scaffold Decision (2026-09-06)

What

End-to-end workbench for training a 0.1B Chinese LLM from scratch. Located: ~/.openclaw/workspace/tiny-llm-studio/

Architecture decisions

  • Llama-style decoder-only: RoPE + RMSNorm + SwiGLU (not GPT-2 style)
  • 0.1B params: 12 layers, 768 hidden, 12 heads, 2048 intermediate → ~108M params
  • Weight tying: embedding = lm_head
  • Pretrain-first: only pretrain and tokenize commands implemented; SFT/DPO/serve deferred
  • Pre-tokenized .bin format: uint16 numpy memmap for fast data loading
  • SentencePiece BPE tokenizer with byte fallback, 32k vocab

Stack

  • PyTorch (MPS + CUDA), SentencePiece, Typer + Rich CLI, Pydantic config, DuckDB (reserved for corpus analytics)
  • Mixed precision (bfloat16 default), gradient accumulation, cosine LR with warmup

Files

tls.yaml          — experiment config
tls/__init__.py
tls/cli.py        — Typer CLI (pretrain, tokenize, info, ui)
tls/config.py     — Pydantic models
tls/model.py      — TinyLlama model (~100 lines)
tls/data.py       — PretrainDataset (memmap .bin reader)
tls/train.py      — training loop
tls/tokenize.py   — SentencePiece train + encode
tls/server.py     — FastAPI + WebSocket bridge to web UI
web/              — React + Vite + Tailwind v4 frontend (7 stages)
web/dist/         — pre-built static assets (served by tls ui)
web/src/lib/bridge.ts    — API client, upgrades sandbox → live
web/src/lib/useBackend.ts — React hook for backend connection state

Web UI (from workspace.tar)

Full simulation workbench with 7 stages: 01 语料数据, 02 分词器, 03 模型架构, 04 训练监控, 05 评测中心, 06 推理沙盒, 07 导出部署. Dark industrial theme (ember/mint/wave palette). Sandbox mode works offline; when tls ui runs, frontend connects via WebSocket for real training metrics.

Wiring

  • tls ui starts FastAPI on :8142, serves web/dist/ + REST/WS APIs
  • Frontend auto-detects backend: shows “沙箱模式” or “后端已连接”
  • WS /ws/train streams real training steps (loss, LR, grad_norm, tok/s)
  • Vite dev proxy forwards /api and /ws to :8142 for hot-reload development

Next steps

  1. Get Chinese corpus (Wikipedia-zh dump, ~2-5GB raw text)
  2. tls tokenize → train tokenizer + encode to .bin
  3. tls pretrain on System76 (CUDA) with real data
  4. cd web && npm run build to rebuild dist after frontend changes
  5. Later: SFT, DPO commands; wire real inference to Stage 06