Skip to content
View sidmanale643's full-sized avatar
💢
💢

Highlights

  • Pro

Block or report sidmanale643

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
sidmanale643/README.md

Sidhant

AI Engineer

Twitter · LinkedIn · GitHub · Email


Stack

Technologies
Languages Python · TypeScript · SQL · Bash
AI / ML PyTorch · CUDA · Transformers · Hugging Face · LangChain · LangGraph · LlamaIndex · MCP
Backend & Data FastAPI · Pydantic · SQLite · PostgreSQL · Redis · LanceDB
Developer Tools Docker · Kubernetes · GCP · Git · GitHub · Supabase · Langfuse

Working On

  • Inference engineering — LLM serving, KV caching, batching, GPU memory management, and runtime performance.
  • AI agents — Stateful coding agents, subagents, tool execution, orchestration, and self-improving harnesses.
  • Memory — Local-first semantic memory, knowledge graphs, hybrid retrieval, and memory decay.
  • Observability & evals — Tracing, datasets, evaluators, regression testing, and CI/CD gates for AI systems.

Projects

  • Helios — Lightweight LLM inference engine built from scratch in PyTorch, with paged attention, prefix caching and continuous batching.
  • Ares — Self improving RLM harness with a persistent IPython environment and recursive subagents.
  • Terminus CLI — Coding harness with its own tool-calling loop, context compaction, and Scout/Worker/Verifier agents coordinated by Mission Control.
  • Atlas — Local-first semantic memory for agents, combining a SQLite knowledge graph, hybrid BM25 + vector search, and exponential memory decay.
  • Evalon — Tracing and evals for Python agents, with datasets, LLM-as-judge, regression checks and CI gates.

Tools for agent users

  • Rewind View — npx rewind-view. Local dashboard that indexes your Claude, Codex, Cursor, OpenCode and Antigravity sessions for search, cost tracking and cross-agent handoffs.
  • CurseBench — uvx cursebench. Measures how often you swear at your coding agents, compared across harnesses and models. Read-only, and nothing leaves your machine.
  • Bloom — Agent-run Obsidian vault following Karpathy's LLM-wiki pattern: drop in papers, videos and articles, and an agent compiles them into a linked knowledge base.

From scratch

  • Qwen3-0.6B from scratch — Qwen3 architecture reimplemented in PyTorch and loaded with the official weights.
  • Inference Engineering — Small, runnable implementations of inference concepts (attention, KV caching and more), built while working toward Helios.

Profile views

Pinned Loading

  1. helios helios Public

    A small, readable PyTorch inference engine for Qwen3-4B with continuous batching, native paged attention, prefix caching, and an OpenAI-compatible chat completions API.

    Python 18 2

  2. terminus-cli terminus-cli Public

    Terminus is an AI coding agent that works directly inside your terminal and understands your codebase. It investigates bugs, builds features, edits code, runs tests, and verifies the result.

    Python 16 1

  3. Ares Ares Public

    Self Improving RLM based coding agent

    TypeScript 5

  4. evalon evalon Public

    Local observability and evaluation for Python agents: trace runs in SQLite, build versioned datasets, and run deterministic or rubric-based evaluations.

    Python 1 1

  5. Atlas Atlas Public

    A Python semantic-memory framework with SQLite storage, optional vector search and local embeddings, plus optional CLI, API, MCP, and web adapters.

    Python 1

  6. atlas-jev atlas-jev Public

    Python 2