A language model built from the ground up with no dependencies on external models.
NeuralForge is a GPT-style decoder-only transformer language model implemented entirely from scratch. No pre-trained weights, no external model dependencies - pure PyTorch from random initialization.
- Pure implementation: No dependencies on existing model weights or architectures
- Scalable: From ~2M to 16B+ parameters
- Modern architecture: Rotary position embeddings (RoPE), SwiGLU feed-forward, RMSNorm
- Dual tokenizers: BPE tokenizer or fast character-level tokenizer
- Flash Attention: SDPA
is_causalfast path for faster training - torch.compile: JIT-compiled training by default (falls back to eager)
- Rich sampling: temperature, top-k, top-p (nucleus), and repetition penalty
- Visual dashboard: Real-time training metrics, GPU stats, loss trends
- Auto validation split: Holds out a slice of the data when none is given
- Named model artifacts: train to
<name>_train.pt, publish<name>.pt, keep<name>_best.pt - GPU-only: CUDA required for training
- KV-cache: Efficient autoregressive text generation
📖 Full Usage Guide → docs/USAGE.md — complete argument
reference for train.py and generate.py, runnable examples, sampling guide,
recommended settings for 12 GB GPUs, best-use-case recipes, and troubleshooting.
A professional, ChatGPT-style web interface with a side-by-side layout: chat with your model on one side, train/track/tune it on the other — all live.
pip install -r requirements.txt # installs fastapi + uvicorn
python -m webui.server # then open http://127.0.0.1:8000- Chat panel — pick a checkpoint and generate, with token-by-token streaming and live sampling controls (temperature, top-k, top-p, repetition penalty).
- Admin panel — configure and launch training (preset, data, tokenizer, epochs, seq-len, batch, lr…), then watch live loss curve, epoch/batch progress, tokens/sec, ETA, best validation loss, and GPU utilization / memory / temperature. Stop a run any time.
Teach the model during a conversation. Every reaction you give — approve, reject, or a corrected answer — becomes a few gradient steps that nudge the model toward what you want, immediately. No big dataset, no offline run needed.
CLI
# Interactive teaching chat (GPU, small model)
python learn.py --checkpoint checkpoints/small.pt --interactive
# Scripted self-test: proves the model learns from basic human interactions
python learn.py --checkpoint checkpoints/small.pt --testIn the interactive session, after each reply type y (approve), n (reject,
then optionally a fix), <text> (teach that better answer), or s (skip).
All interactions are logged to checkpoints/feedback.jsonl and can be exported
with --export data/learned_interactions.txt for a full offline training run.
Web UI
In NeuralForge Studio, each assistant reply has a feedback toolbar (👍 approve, 👎 reject, ✏ correct). Reacting runs live gradient steps on the selected model; the Live Learning card shows interactions/steps and last loss, and can Save the learned weights or Export the corpus.
See neuralforge/learning/online.py for the OnlineLearner implementation
(teach / approve / reject primitives, masked next-token loss, JSONL store).
git clone https://github.com/UDAIE-A/NeuralForge.git
cd NeuralForge
python -m venv venv
.\venv\Scripts\Activate.ps1
pip install torch --index-url https://download.pytorch.org/whl/cu126Put your training data in a text file:
# Any text file works - books, articles, code, etc.
# Larger data = better results# Fast training with character-level tokenizer (recommended for quick tests)
python train.py --preset small --data data/train.txt --epochs 100 --batch-size 64 --seq-len 512 --char
# Better quality with BPE tokenizer (slower tokenizer training)
python train.py --preset small --data data/train.txt --epochs 100 --batch-size 64 --seq-len 512python generate.py --checkpoint checkpoints/small.pt --prompt "Alice" --max-tokens 200
# Better quality sampling: nucleus sampling + repetition penalty
python generate.py --checkpoint checkpoints/small.pt --prompt "Alice" \
--max-tokens 200 --top-p 0.9 --repetition-penalty 1.2
# Interactive mode
python generate.py --checkpoint checkpoints/small.pt --interactiveParameter counts use each preset's default vocab size (the embedding scales with the actual tokenizer vocab at train time).
| Preset | Parameters | d_model | n_heads | n_layers | d_ff | VRAM (approx) |
|---|---|---|---|---|---|---|
| tiny | ~2M | 128 | 4 | 4 | 512 | ~2 GB |
| small | ~12M | 256 | 8 | 8 | 1024 | ~4 GB |
| base | ~138M | 768 | 12 | 12 | 3072 | ~8 GB |
| large | ~435M | 1024 | 16 | 24 | 4096 | ~12 GB |
| xl | ~2.2B | 2048 | 32 | 32 | 8192 | ~24 GB |
| xxl | ~22B | 4096 | 32 | 80 | 16384 | multi-GPU |
- Instant training (no tokenizer learning needed)
- Faster training on small datasets
- Lower quality output
- Good for quick experiments
- Learns subword units from data
- Better quality output
- Slower tokenizer training
- Recommended for serious training
neuralforge/
├── core/
│ ├── config.py # Model configuration (tiny to xxl)
│ └── model.py # Transformer: RoPE, SwiGLU, RMSNorm, Flash Attention
├── tokenizer/
│ ├── bpe.py # BPE tokenizer from scratch
│ └── char_tokenizer.py # Character-level tokenizer
├── training/
│ ├── data.py # Dataset and DataLoader
│ └── trainer.py # Training with visual dashboard
├── train.py # Main training script
├── generate.py # Text generation
├── data/ # Training data
└── checkpoints/ # Saved models
Real-time metrics during training:
- Epoch and batch progress bars
- Loss value and trend sparkline
- GPU utilization, memory, temperature
- Tokens per second throughput
- ETA and elapsed time
- Ctrl+C saves checkpoint and shows resume command
- Python 3.8+
- PyTorch 2.0+
- NVIDIA GPU with CUDA support
- 4-12 GB VRAM depending on model size
- Transformer architecture from scratch
- BPE tokenizer
- Character-level tokenizer
- Flash Attention
- Rotary position embeddings (RoPE)
- SwiGLU feed-forward + RMSNorm
- top-p / repetition-penalty sampling
- torch.compile training
- Visual training dashboard
- Live online fine-tuning (NeuralForge Learn)
- GPU-only training
- Multi-GPU training
- Gradient checkpointing for large models
- Mixture of Experts for scaling
- RLHF alignment
- Instruction tuning
The model architecture changed (RoPE, SwiGLU, RMSNorm), so checkpoints
trained before that switch cannot be loaded by the current code. The previous
architecture is preserved at the v0-legacy-arch tag:
# Generate from old (pre-RoPE) checkpoints
git checkout v0-legacy-arch
# Return to the current architecture
git checkout mainNew checkpoints trained on main are the way forward and should produce
better results.
MIT License - see LICENSE for details.
UDAIE-A - GitHub