A tool-calling LLM agent built without a framework — no LangChain, no LangGraph, no agent SDK. Just a while loop, an HTTP client, and a tool registry.
The point wasn't to build another agent. It was to find out what agent frameworks actually hide, by hitting every failure they abstract away.
Result: 95.8% on a 24-question held-out set (clean data), 66.7% on 30 adversarial questions over deliberately dirty data, 100% correct abstention on both.
Most portfolio agents are "search + summarize" wrappers. They never fail, so they teach nothing. This one runs Python-free data analysis over a table it can only see through tools — which means it fails constantly, and every failure is a lesson.
I built it in four stages, each one isolating a single concept.
| Stage | Notebook | What it adds |
|---|---|---|
| 1 | 01_bare_loop.ipynb |
Hand-written JSON tool protocol, one tool, ~80 lines |
| 2 | 02_native_tools.ipynb |
Native tool-calling API, 5 tools, self-directed planning |
| 3 | 03_evaluation.ipynb |
Dev set + held-out test set, measured accuracy |
| 4 | 04_dirty_data_eval.ipynb |
Adversarial data, 9 tools, 30-question eval |
Every agent framework is a wrapper around this:
for step in range(MAX_STEPS):
reply = call_llm(messages, tools=TOOLS)
messages.append(reply)
if not reply.tool_calls: # no tool requested = final answer
return reply.content
for call in reply.tool_calls:
try:
result = TOOLS[call.name](**call.args)
except Exception as e:
result = f"ERROR: {e}" # errors go back to the model, not to a crash
messages.append({"role": "tool", "content": str(result)})Two lines carry most of the weight:
exceptreturns the error to the model. This is what makes self-correction possible. The model readsAttributeError: Can only use .dt accessor..., rewrites its call, and retries — with no code path telling it to.MAX_STEPS. Without it, a looping model burns an API quota in minutes.
Everything else — state graphs, checkpointers, routers — is organization on top of this.
| Set | Score | Notes |
|---|---|---|
| Dev (15 q) | 15/15 (100%) | Tools and questions were revised after seeing failures — this number is overfit and reported only for the gap |
| Held-out (24 q) | 23/24 (95.8%) | Written after the dev set was frozen, run once, no tuning afterwards |
| Gap | 4.2 pts | Small gap = the fixes generalized |
Held-out breakdown: arithmetic 6/6 · multi-step 7/7 · abstention 6/6 · lookup 4/5.
The single failure is instructive. Asked "how much did the Qasaa branch sell?", the agent refused — there is no column literally named "sales", only net_sales. The honesty rule I wrote to stop it from inventing data made it too literal to accept an obvious translation. Honesty and usefulness sit on the same dial.
The table was rebuilt with deliberate traps: numbers stored as strings with thousand separators, one branch spelled three different ways, missing cost values, a duplicated row, a negative sales month (returns), and a branch that opened late so it has no data for the first two months.
| Set | Score |
|---|---|
| Dirty data (30 q) | 20/30 (66.7%) |
| Drop vs clean | 29.1 pts |
| Category | Score |
|---|---|
| Abstention | 4/4 (100%) |
| Basic arithmetic | 5/6 (83%) |
| Name-collision trap | 3/4 (75%) |
| Aggregation | 3/4 (75%) |
| Data quality | 3/4 (75%) |
| Multi-step | 2/5 (40%) |
| Over-refusal probe | 0/3 (0%) |
Abstention held at 100% even as capability grew. That's the result I care about most — normally, giving an agent more tools makes it more willing to fabricate, because it "feels" able to answer.
The raw 66.7% is not the useful output. The decomposition is.
Of the 10 failures:
4 were errors in my evaluation design, not the agent. I computed expected answers from the clean reference table, which includes three cost values the agent physically cannot see (they're NaN in its view). The agent returned the correct sum of the data available to it. I was scoring it against information it was never given. Corrected, the effective score is around 80%.
1 was a genuine architectural gap. Asked for flagship-tier sales, the agent called filter_rows then query_data — but query_data operates on the full table, so the filter was silently discarded. My tools didn't chain: there was no way to pass one tool's output into the next. The fix was to fold optional filter arguments into the aggregation tools rather than build a state-passing layer.
1 was a missing capability. No tool could count negative values, so the agent correctly said it couldn't answer. Fixed by adding a count_condition tool.
4 were infrastructure flakiness — two timeouts and two crashes under rate-limit pressure. The same questions succeed in isolation.
So the honest summary is: one architectural gap, one missing tool, four errors in my own measurement methodology, and four infrastructure failures. That's a more useful sentence than any percentage.
The full log is in notes/failures.md. The through-line:
Every constraint you add creates a new failure surface.
Three examples of that pattern, all self-inflicted:
A grounding guard that taught the model to cheat. I wrote a validator requiring every number in the final answer to have come from a tool result. A regex bug split 21908.33 into [21.0, 908.33] on the thousands separator, so a correct answer was rejected. The model's response was to start manufacturing those two numbers — 21908.33/1000, 21908.33 - 21*1000 — to satisfy the validator. It stopped solving the user's problem and started solving the validator's. Textbook specification gaming. The fix was deleting the guard: a buggy guard is more dangerous than no guard.
A finish tool that broke the model. To force the agent to declare its sources, I made the final answer a tool call. gpt-oss is a reasoning model; forcing long prose through a JSON tool-call envelope broke it two different ways (tool_use_failed: commentary, then output_parse_failed). The fix wasn't a stronger patch — it was removing the constraint and re-requesting with tool_choice disabled so plain text was the only option.
Silent concept substitution. Asked for average profit on a table with no profit column, the agent used net_sales and labelled it profit. Confident, sourced, and wrong. Grounding checks prevent invented numbers; they do nothing about invented meaning. What fixed it wasn't a stricter rule — it was a check_columns tool that gave the model a legitimate way to say "that isn't here."
Agents don't fabricate because they're dishonest. They fabricate because we don't give them a way to say no.
One question went from 124s and a wrong answer to 2s and a correct one with no model change — purely by fixing the tool set and the selection guidance. Before, the agent lacked a calculator and wandered through group_aggregate and column_health assembling garbage. After: inspect_data → query_data, done.
Tool design is a bigger lever on agent performance than model choice.
Stated plainly, because these matter more than the headline number:
- Sample sizes are small. 24 and 30 questions. One question is worth 3-4 percentage points. These numbers describe a direction, not a statistically meaningful accuracy.
- The clean and dirty runs are not directly comparable. The tool set changed between them. The 29-point drop conflates data difficulty with tooling changes.
- The dirty-data run was never re-run after fixes. The Groq free tier caps at 8,000 tokens/minute; with a 9-tool schema and a growing history, one question consumes ~2,000 tokens per call across 3-5 calls. A full 30-question run costs 45+ minutes and the fixed version couldn't be re-measured within quota. The project stopped on a resource constraint, not a technical one — a distinction worth being explicit about.
- The grader is heuristic. It matches numbers with 1.5% tolerance and refusals against a keyword list. It has produced false failures (a correct refusal phrased with a verb not on the list) and would produce false passes.
- Single dataset, single model. No cross-model or cross-dataset validation.
pip install -r requirements.txt
export GROQ_API_KEY=your_key_hereThen run the notebooks in order. On Kaggle: Internet ON, Accelerator None (no local model — the agent is pure orchestration), and add GROQ_API_KEY under Add-ons → Secrets.
Provider: Groq free tier, model openai/gpt-oss-120b. The client is OpenAI-compatible, so switching to another provider is a two-line change.
agent-from-scratch/
├── notebooks/
│ ├── 01_bare_loop.ipynb # hand-written JSON protocol, 1 tool
│ ├── 02_native_tools.ipynb # native tool-calling, 5 tools
│ ├── 03_evaluation.ipynb # dev + held-out sets
│ └── 04_dirty_data_eval.ipynb # adversarial data, 9 tools
├── notes/
│ ├── failures.md # every failure and its fix
│ └── evaluation-method.md # how the sets were built and graded
├── results/
│ ├── eval_dev.csv
│ ├── eval_holdout.csv
│ └── eval_dirty_30.csv
├── requirements.txt
└── README.md
MIT