AI Engineer
Twitter · LinkedIn · GitHub · Email
| Technologies | |
|---|---|
| Languages | Python · TypeScript · SQL · Bash |
| AI / ML | PyTorch · CUDA · Transformers · Hugging Face · LangChain · LangGraph · LlamaIndex · MCP |
| Backend & Data | FastAPI · Pydantic · SQLite · PostgreSQL · Redis · LanceDB |
| Developer Tools | Docker · Kubernetes · GCP · Git · GitHub · Supabase · Langfuse |
- Inference engineering — LLM serving, KV caching, batching, GPU memory management, and runtime performance.
- AI agents — Stateful coding agents, subagents, tool execution, orchestration, and self-improving harnesses.
- Memory — Local-first semantic memory, knowledge graphs, hybrid retrieval, and memory decay.
- Observability & evals — Tracing, datasets, evaluators, regression testing, and CI/CD gates for AI systems.
- Helios — Lightweight LLM inference engine built from scratch in PyTorch, with paged attention, prefix caching and continuous batching.
- Ares — Self improving RLM harness with a persistent IPython environment and recursive subagents.
- Terminus CLI — Coding harness with its own tool-calling loop, context compaction, and Scout/Worker/Verifier agents coordinated by Mission Control.
- Atlas — Local-first semantic memory for agents, combining a SQLite knowledge graph, hybrid BM25 + vector search, and exponential memory decay.
- Evalon — Tracing and evals for Python agents, with datasets, LLM-as-judge, regression checks and CI gates.
- Rewind View —
npx rewind-view. Local dashboard that indexes your Claude, Codex, Cursor, OpenCode and Antigravity sessions for search, cost tracking and cross-agent handoffs. - CurseBench —
uvx cursebench. Measures how often you swear at your coding agents, compared across harnesses and models. Read-only, and nothing leaves your machine. - Bloom — Agent-run Obsidian vault following Karpathy's LLM-wiki pattern: drop in papers, videos and articles, and an agent compiles them into a linked knowledge base.
- Qwen3-0.6B from scratch — Qwen3 architecture reimplemented in PyTorch and loaded with the official weights.
- Inference Engineering — Small, runnable implementations of inference concepts (attention, KV caching and more), built while working toward Helios.



