From a single NAND gate to a 16-bit ALU
Complete source code for the companion video, showing how a single NAND gate can calculate 7 + 8 = 15
This repository contains the complete source code for the companion video. It builds up from a single NAND gate to a 16-bit ALU.
Episode Goal: Understand how a single logic gate (NAND) can eventually calculate 7 + 8 = 15 through progressive building blocks.
- NAND-only Implementation: Every component built from single NAND gate primitive
- Progressive Building: From logic gates → half adder → full adder → 8-bit ALU → 16-bit ALU
- Complete Toolchain: Verilog RTL, testbenches, Python assembler, and build automation
- Educational Focus: Clear progression following the video episode structure
- Hands-on Validation: Live testing with calculator verification (7 + 8 = 15)
- Open Source: All code and documentation freely available on GitHub
nand2cpu/ # 16-bit ALU from NAND gates
├── src/ # Source code
│ ├── rtl/ # Verilog RTL modules
│ │ ├── nand_gate.v # Universal NAND gate primitive
│ │ ├── and_gate.v # AND gate (2× NAND)
│ │ ├── or_gate.v # OR gate (3× NAND)
│ │ ├── alu8.v # 8-bit ALU (7 operations)
│ │ └── alu16.v # 16-bit ALU (chained from 8-bit)
│ ├── fpga/ # FPGA implementation files
│ │ ├── top.v # Top-level FPGA module
│ │ └── top.xdc # Timing constraints
│ └── testbenches/ # Simulation testbenches
│ ├── tb_nand_gate.v # NAND gate verification
│ ├── tb_alu16.v # 16-bit ALU verification
│ └── add7_plus_8.v # 7+8=15 demonstration
├── tools/ # Episode development tools
│ ├── assembler/ # Python assembler (featured in video)
│ │ ├── main.py # Main assembler interface
│ │ ├── parser.py # Instruction parser
│ │ └── encoder.py # 16-bit machine code generator
│ └── scripts/ # Build automation
│ ├── build_sim.sh # Simulation build
│ └── build_fpga.tcl # FPGA synthesis
├── docs/ # Episode documentation
│ ├── Script ENG.md # Complete video script
│ └── Slides ENG.md # Presentation slides
├── examples/ # Assembly examples
│ └── test_vectors.asm # Sample assembly code
└── Makefile # Build automation (make help)
# Ubuntu/Debian
sudo apt update
sudo apt install iverilog python3 make
# macOS
brew install icarus-verilog python3# Clone repository
git clone https://github.com/promaaa/nand2cpu.git
cd nand2cpu
# Reproduce the video demonstration: 7 + 8 = 15
make sim-add7_plus_8
# Test the complete 16-bit ALU
make sim-tb_alu16
# Test individual components
make sim-tb_nand_gate # Test NAND gate primitive
# Test the Python assembler (featured in episode)
make assembler
hexdump -C tools/assembler/test.bin| Component | Description | NAND Gates | Episode Section |
|---|---|---|---|
nand_gate.v |
Universal primitive | 1 | Foundation |
and_gate.v |
4-bit AND gate | 2 | Basic Logic |
or_gate.v |
OR gate | 3 | Logic Gates |
alu8.v |
8-bit ALU | ~200 | Main Build |
alu16.v |
16-bit ALU | ~400 | Scaling Up |
This repository implements the exact demonstration from the video:
// From add7_plus_8.v testbench
initial begin
A = 8'b00000111; // 7 in binary
B = 8'b00001000; // 8 in binary
op = 3'b000; // ADD operation
#10;
// Expected: Result = 15 (0b00001111)
$display("7 + 8 = %d", Result);
endVisual confirmation: The simulator shows the exact same result as a calculator!
8-bit ALU (alu8.v) - Core of the episode
- Operations: ADD, SUB, AND, OR, XOR, SHL, SHR, NOT
- Construction: Half-adder → Full-adder → 8-bit chain
- Flags: Zero, Carry, Overflow detection
- Latency: 0.8ns (as shown in video benchmarks)
16-bit ALU (alu16.v) - Scaling demonstration
- Architecture: Two chained 8-bit ALUs
- Carry propagation: Between upper and lower bytes
- Latency: 1.6ns (2× scaling challenge discussed)
- Pipelining: Concepts introduced for performance optimization
The Python assembler demonstrated in the episode converts assembly instructions into 16-bit machine code, exactly as shown in the video.
# Featured in video: 3-operand instructions
ADD R0, R1, R2 ; R0 = R1 + R2
SUB R3, R4, R5 ; R3 = R4 - R5
AND R6, R7, R0 ; R6 = R7 & R0
# 2-operand instructions
SHL R1, R2 ; R1 = R2 << 1
NOT R3, R4 ; R3 = ~R4[15:13] [12:10] [9:7] [6:4] [3:0]
Opcode Rd Rs1 Rs2 Unused
Parser: Tokenizes mnemonics and operands, validates syntax
Encoder: Maps to 4-bit opcodes, packs into 16-bit words
# Reproduce the exact video demonstration
make sim-add7_plus_8 # 7 + 8 = 15 calculation
# Validate NAND gate foundation
make sim-tb_nand_gate # Truth table verification
# Test complete 16-bit ALU
make sim-tb_alu16 # All 7 operations + carry propagation
# Verify assembler functionality
make assembler # Parser + Encoder testing# Quick episode validation
make quick-test # NAND + ALU core tests
make validate # Repository structure check
# Full test suite
make test-gates # All logic gate tests
make test-alu # Complete ALU validation
make test-all # Everything (as shown in episode)| Test | Episode Focus | Status |
|---|---|---|
add7_plus_8.v |
Main demonstration | ✅ |
tb_nand_gate.v |
Foundation primitive | ✅ |
tb_alu16.v |
16-bit scaling | ✅ |
| Python Assembler | Tool demonstration | ✅ |
Coming in Part 2: Neural Networks on Microcontrollers
- Model compression and quantization
- Real-time inference on STM32
- Energy efficiency analysis
- Live Edge AI demonstrations
Full Series Roadmap:
- Part 3: Memory Systems and Cache Hierarchy
- Part 4: Complete CPU Architecture
- Part 5: FPGA Implementation and Synthesis
- Part 6: Custom AI Accelerator Hardware
Found this episode helpful? Contributions welcome!
Episode-specific improvements:
- Additional test cases for the 7+8=15 demonstration
- Alternative ALU implementations
- Extended assembler instruction set
- Performance optimizations
Documentation:
- Code comments and explanations
- Additional examples following video structure
- Educational enhancements
- Fork the repository
- Create feature branch (
git checkout -b feature/episode-improvement) - Commit changes (
git commit -m 'Enhance episode demonstration') - Push to branch (
git push origin feature/episode-improvement) - Open Pull Request
This project is part of the "From Bits to Chip" educational series.
Distributed under the MIT License - see LICENSE for details.
- Nand2Tetris Course: Educational methodology and inspiration for bottom-up approach
- MIT 6.004: Computer architecture foundations demonstrated in this episode
- Hardware Description Community: Verilog best practices and simulation techniques
Enjoyed Part 1? Please star ⭐️ this repository!
Subscribe to the series: YouTube Channel • Follow on GitHub
Educational series: From Bits to Chip - Part 1 of 6