DeepSeek Harness benchmark path: run, validate, report #62
initial-d
announced in
Announcements
Replies: 1 comment
|
Follow-up: I added a broader quant-agent framing page for teams evaluating LLM trading agents or agent harnesses: https://github.com/initial-d/ml-quant-trading/blob/main/docs/quant_agent_reproducibility_target.md The point is to offer a fixed reproducibility target: run the documented benchmark, preserve I also opened a listing PR to Awesome LLM Quantitative Trading Papers under the financial benchmark/evaluation angle: |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
A small DeepSeek Harness path is now available for
ml-quant-tradingbenchmark reports.The optional plugin lives here:
https://github.com/initial-d/dsh-plugin-mlquant-benchmark
It is intentionally narrow. It gives a DSH agent four tools:
artifacts/benchmark-v1.json;Install:
Suggested prompt:
Reports are welcome through the dedicated template:
https://github.com/initial-d/ml-quant-trading/issues/new?template=deepseek_harness_benchmark.yml
The useful question is simple: can a coding agent reproduce the benchmark, preserve the artifact, and avoid overstating runtime numbers as trading evidence?
All reactions