4 GPU workers · 4 models born here

Train your first
decoder-only model
from scratch.

Not a fine-tune. Not an adapter. Random weights, a real corpus, and a real loss curve you watch fall in real time — from a browser tab, on our GPUs.

Signing in opens NanoDex full screen — Hugging Face's login page can't run inside an embedded frame.

1

Pick a size

Four architectures, from 492k to 8.1M parameters. That's the real total — embeddings included, no asterisk.

2

Pick a token budget

Anywhere from 1M to 1.5 billion tokens of fineweb-edu. More tokens, sharper model, longer wait.

3

Watch it learn

Your run joins the queue. Live loss curve, throughput, logs — then the weights land in your own Hugging Face account.

The four sizes

Small enough to finish. Big enough to learn something.

NanoDex-100K
100,064 params
2 layers · 32 hidden
2 heads (1 KV) · FFN 147
NanoDex-250K
250,224 params
4 layers · 48 hidden
4 heads (2 KV) · FFN 215
NanoDex-500K
499,600 params
4 layers · 80 hidden
5 heads (1 KV) · FFN 285
NanoDex-1M
1,000,224 params
5 layers · 96 hidden
6 heads (2 KV) · FFN 472
NanoDex-2M
2,000,256 params
6 layers · 128 hidden
8 heads (2 KV) · FFN 647
NanoDex-5M
4,999,872 params
8 layers · 192 hidden
8 heads (2 KV) · FFN 839
NanoDex-10M
10,001,664 params
10 layers · 256 hidden
8 heads (2 KV) · FFN 1020
NanoDex-25M
24,993,600 params
12 layers · 320 hidden
10 heads (2 KV) · FFN 1856
NanoDex-50M
49,995,456 params
12 layers · 448 hidden
8 heads (2 KV) · FFN 2669
NanoDex-100M
95,095,680 params
21 layers · 640 hidden
8 heads (2 KV) · FFN 1792
GPT-2 Small
114,838,272 params
12 layers · 768 hidden
12 heads (12 KV) · FFN 3072
What you actually get

A model that is genuinely yours

Every finished run is pushed to your Hugging Face namespace: config.json, model.safetensors, the tokenizer, a training_run.json with the full recipe, and a generated model card. Load it with transformers like anything else on the Hub.

Honest expectations

It will not be an assistant

At a few hundred thousand parameters, a model learns word shapes, common collocations and a little syntax. It will produce English-looking text with no facts in it. That's the correct outcome — and watching cross-entropy fall from 7.6 (uniform noise over 2,048 tokens) to somewhere near 4 is the whole point.

Trained here

The newest models on the shelf

Explore all →

Ready when you are.

Sign in with Hugging Face — no separate account, no password.

Start a run  →