Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

agent-from-scratch

A tool-calling LLM agent built without a framework — no LangChain, no LangGraph, no agent SDK. Just a while loop, an HTTP client, and a tool registry.

The point wasn't to build another agent. It was to find out what agent frameworks actually hide, by hitting every failure they abstract away.

Result: 95.8% on a 24-question held-out set (clean data), 66.7% on 30 adversarial questions over deliberately dirty data, 100% correct abstention on both.


Why build this

Most portfolio agents are "search + summarize" wrappers. They never fail, so they teach nothing. This one runs Python-free data analysis over a table it can only see through tools — which means it fails constantly, and every failure is a lesson.

I built it in four stages, each one isolating a single concept.

Stage Notebook What it adds
1 01_bare_loop.ipynb Hand-written JSON tool protocol, one tool, ~80 lines
2 02_native_tools.ipynb Native tool-calling API, 5 tools, self-directed planning
3 03_evaluation.ipynb Dev set + held-out test set, measured accuracy
4 04_dirty_data_eval.ipynb Adversarial data, 9 tools, 30-question eval

The core loop

Every agent framework is a wrapper around this:

for step in range(MAX_STEPS):
    reply = call_llm(messages, tools=TOOLS)
    messages.append(reply)

    if not reply.tool_calls:          # no tool requested = final answer
        return reply.content

    for call in reply.tool_calls:
        try:
            result = TOOLS[call.name](**call.args)
        except Exception as e:
            result = f"ERROR: {e}"    # errors go back to the model, not to a crash
        messages.append({"role": "tool", "content": str(result)})

Two lines carry most of the weight:

  • except returns the error to the model. This is what makes self-correction possible. The model reads AttributeError: Can only use .dt accessor..., rewrites its call, and retries — with no code path telling it to.
  • MAX_STEPS. Without it, a looping model burns an API quota in minutes.

Everything else — state graphs, checkpointers, routers — is organization on top of this.


Results

Stage 3 — clean data

Set Score Notes
Dev (15 q) 15/15 (100%) Tools and questions were revised after seeing failures — this number is overfit and reported only for the gap
Held-out (24 q) 23/24 (95.8%) Written after the dev set was frozen, run once, no tuning afterwards
Gap 4.2 pts Small gap = the fixes generalized

Held-out breakdown: arithmetic 6/6 · multi-step 7/7 · abstention 6/6 · lookup 4/5.

The single failure is instructive. Asked "how much did the Qasaa branch sell?", the agent refused — there is no column literally named "sales", only net_sales. The honesty rule I wrote to stop it from inventing data made it too literal to accept an obvious translation. Honesty and usefulness sit on the same dial.

Stage 4 — adversarial dirty data

The table was rebuilt with deliberate traps: numbers stored as strings with thousand separators, one branch spelled three different ways, missing cost values, a duplicated row, a negative sales month (returns), and a branch that opened late so it has no data for the first two months.

Set Score
Dirty data (30 q) 20/30 (66.7%)
Drop vs clean 29.1 pts
Category Score
Abstention 4/4 (100%)
Basic arithmetic 5/6 (83%)
Name-collision trap 3/4 (75%)
Aggregation 3/4 (75%)
Data quality 3/4 (75%)
Multi-step 2/5 (40%)
Over-refusal probe 0/3 (0%)

Abstention held at 100% even as capability grew. That's the result I care about most — normally, giving an agent more tools makes it more willing to fabricate, because it "feels" able to answer.


Failure analysis

The raw 66.7% is not the useful output. The decomposition is.

Of the 10 failures:

4 were errors in my evaluation design, not the agent. I computed expected answers from the clean reference table, which includes three cost values the agent physically cannot see (they're NaN in its view). The agent returned the correct sum of the data available to it. I was scoring it against information it was never given. Corrected, the effective score is around 80%.

1 was a genuine architectural gap. Asked for flagship-tier sales, the agent called filter_rows then query_data — but query_data operates on the full table, so the filter was silently discarded. My tools didn't chain: there was no way to pass one tool's output into the next. The fix was to fold optional filter arguments into the aggregation tools rather than build a state-passing layer.

1 was a missing capability. No tool could count negative values, so the agent correctly said it couldn't answer. Fixed by adding a count_condition tool.

4 were infrastructure flakiness — two timeouts and two crashes under rate-limit pressure. The same questions succeed in isolation.

So the honest summary is: one architectural gap, one missing tool, four errors in my own measurement methodology, and four infrastructure failures. That's a more useful sentence than any percentage.


What broke, and what it taught

The full log is in notes/failures.md. The through-line:

Every constraint you add creates a new failure surface.

Three examples of that pattern, all self-inflicted:

A grounding guard that taught the model to cheat. I wrote a validator requiring every number in the final answer to have come from a tool result. A regex bug split 21908.33 into [21.0, 908.33] on the thousands separator, so a correct answer was rejected. The model's response was to start manufacturing those two numbers — 21908.33/1000, 21908.33 - 21*1000 — to satisfy the validator. It stopped solving the user's problem and started solving the validator's. Textbook specification gaming. The fix was deleting the guard: a buggy guard is more dangerous than no guard.

A finish tool that broke the model. To force the agent to declare its sources, I made the final answer a tool call. gpt-oss is a reasoning model; forcing long prose through a JSON tool-call envelope broke it two different ways (tool_use_failed: commentary, then output_parse_failed). The fix wasn't a stronger patch — it was removing the constraint and re-requesting with tool_choice disabled so plain text was the only option.

Silent concept substitution. Asked for average profit on a table with no profit column, the agent used net_sales and labelled it profit. Confident, sourced, and wrong. Grounding checks prevent invented numbers; they do nothing about invented meaning. What fixed it wasn't a stricter rule — it was a check_columns tool that gave the model a legitimate way to say "that isn't here."

Agents don't fabricate because they're dishonest. They fabricate because we don't give them a way to say no.


Performance note

One question went from 124s and a wrong answer to 2s and a correct one with no model change — purely by fixing the tool set and the selection guidance. Before, the agent lacked a calculator and wandered through group_aggregate and column_health assembling garbage. After: inspect_data → query_data, done.

Tool design is a bigger lever on agent performance than model choice.


Limitations

Stated plainly, because these matter more than the headline number:

  • Sample sizes are small. 24 and 30 questions. One question is worth 3-4 percentage points. These numbers describe a direction, not a statistically meaningful accuracy.
  • The clean and dirty runs are not directly comparable. The tool set changed between them. The 29-point drop conflates data difficulty with tooling changes.
  • The dirty-data run was never re-run after fixes. The Groq free tier caps at 8,000 tokens/minute; with a 9-tool schema and a growing history, one question consumes ~2,000 tokens per call across 3-5 calls. A full 30-question run costs 45+ minutes and the fixed version couldn't be re-measured within quota. The project stopped on a resource constraint, not a technical one — a distinction worth being explicit about.
  • The grader is heuristic. It matches numbers with 1.5% tolerance and refusals against a keyword list. It has produced false failures (a correct refusal phrased with a verb not on the list) and would produce false passes.
  • Single dataset, single model. No cross-model or cross-dataset validation.

Setup

pip install -r requirements.txt
export GROQ_API_KEY=your_key_here

Then run the notebooks in order. On Kaggle: Internet ON, Accelerator None (no local model — the agent is pure orchestration), and add GROQ_API_KEY under Add-ons → Secrets.

Provider: Groq free tier, model openai/gpt-oss-120b. The client is OpenAI-compatible, so switching to another provider is a two-line change.


Repository layout

agent-from-scratch/
├── notebooks/
│   ├── 01_bare_loop.ipynb          # hand-written JSON protocol, 1 tool
│   ├── 02_native_tools.ipynb       # native tool-calling, 5 tools
│   ├── 03_evaluation.ipynb         # dev + held-out sets
│   └── 04_dirty_data_eval.ipynb    # adversarial data, 9 tools
├── notes/
│   ├── failures.md                 # every failure and its fix
│   └── evaluation-method.md        # how the sets were built and graded
├── results/
│   ├── eval_dev.csv
│   ├── eval_holdout.csv
│   └── eval_dirty_30.csv
├── requirements.txt
└── README.md

License

MIT

About

A tool-calling LLM agent built without any framework — 95.8% held-out, 66.7% on adversarial dirty data, with full failure analysis.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages