Tiny LLM Studio — Scaffold Decision (2026-09-06)
What
End-to-end workbench for training a 0.1B Chinese LLM from scratch.
Located: ~/.openclaw/workspace/tiny-llm-studio/
Architecture decisions
- Llama-style decoder-only: RoPE + RMSNorm + SwiGLU (not GPT-2 style)
- 0.1B params: 12 layers, 768 hidden, 12 heads, 2048 intermediate → ~108M params
- Weight tying: embedding = lm_head
- Pretrain-first: only
pretrainandtokenizecommands implemented; SFT/DPO/serve deferred - Pre-tokenized .bin format: uint16 numpy memmap for fast data loading
- SentencePiece BPE tokenizer with byte fallback, 32k vocab
Stack
- PyTorch (MPS + CUDA), SentencePiece, Typer + Rich CLI, Pydantic config, DuckDB (reserved for corpus analytics)
- Mixed precision (bfloat16 default), gradient accumulation, cosine LR with warmup
Files
tls.yaml — experiment config
tls/__init__.py
tls/cli.py — Typer CLI (pretrain, tokenize, info, ui)
tls/config.py — Pydantic models
tls/model.py — TinyLlama model (~100 lines)
tls/data.py — PretrainDataset (memmap .bin reader)
tls/train.py — training loop
tls/tokenize.py — SentencePiece train + encode
tls/server.py — FastAPI + WebSocket bridge to web UI
web/ — React + Vite + Tailwind v4 frontend (7 stages)
web/dist/ — pre-built static assets (served by tls ui)
web/src/lib/bridge.ts — API client, upgrades sandbox → live
web/src/lib/useBackend.ts — React hook for backend connection state
Web UI (from workspace.tar)
Full simulation workbench with 7 stages:
01 语料数据, 02 分词器, 03 模型架构, 04 训练监控,
05 评测中心, 06 推理沙盒, 07 导出部署.
Dark industrial theme (ember/mint/wave palette). Sandbox mode works offline;
when tls ui runs, frontend connects via WebSocket for real training metrics.
Wiring
tls uistarts FastAPI on :8142, serves web/dist/ + REST/WS APIs- Frontend auto-detects backend: shows “沙箱模式” or “后端已连接”
- WS /ws/train streams real training steps (loss, LR, grad_norm, tok/s)
- Vite dev proxy forwards /api and /ws to :8142 for hot-reload development
Next steps
- Get Chinese corpus (Wikipedia-zh dump, ~2-5GB raw text)
tls tokenize→ train tokenizer + encode to .bintls pretrainon System76 (CUDA) with real datacd web && npm run buildto rebuild dist after frontend changes- Later: SFT, DPO commands; wire real inference to Stage 06