Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,79 @@
{
"meta": {
"scenario": "creative-writing",
"mode": "single-shot",
"label": "m5-max-48gb-40gpu ollama",
"backend": "ollama",
"model_info": {
"name": "llama3.1:8b",
"format": "gguf",
"family": "llama",
"parameter_size": "8.0B",
"quantization": "Q4_K_M"
},
"runs": 1,
"max_tokens": 2000,
"cold": false,
"warm_up_time": 0.183,
"system": {
"os": "Darwin",
"os_version": "26.6",
"arch": "arm64",
"chip": "Apple M5 Max",
"memory_gb": 48,
"cpu_cores": 18,
"cpu_cores_performance": 6,
"cpu_cores_efficiency": 12,
"gpu_cores": 40,
"gpu_wired_limit_mb": 0
},
"timestamp": "2026-07-28T15:30:51"
},
"results": [
{
"turn": 1,
"run": 1,
"ctx_tokens_est": 57,
"new_tokens_est": 29,
"ttft": 0.149,
"gen_time": 3.839,
"gen_tps": 88.8,
"total": 3.987,
"output_tokens": 341,
"prompt_eval_count": 67,
"prompt_eval_duration_ms": 58.5,
"_response": "In subterranean silence, they reside,\nA murmuration of machines, side by side.\nTheir soft susurrations weave a gentle hush,\nAs bytes and bits converge, a digital rush.\n\nIn this dimly lit lair, where shadows play,\nThey stand vigilant, night and endless day.\nTheir humming hearts beat in harmony and time,\nPreserving the past, in virtual prime.\n\nA family's history, a tapestry so fine,\nUnfurls within their circuitry's gentle design.\nLaughter, tears, and whispers of love now past,\nAre stored within these servers, forever to last.\n\nPhotographs of holidays, birthday cakes aglow,\nAnd snapshots of childhood, with memories that grow.\nThe creaks of floorboards, beneath the weight of years,\nAre digitized and preserved through joyous tears.\n\nTheir processors hum, a soothing serenade,\nAs mother's recipes, and father's favorite shade,\nOf blue are saved within these digital stores,\nA sanctuary for moments they've lived before.\n\nIn this underground sanctum, where data reigns,\nThe family's legacy is safeguarded from life's pains.\nTheir memories, a testament to love's enduring might,\nA treasure trove of moments, shining with delight.\n\nAs seasons change above, and years go by like sand,\nThese servers remain, steadfast, hand in digital hand.\nTheir hum, a lullaby, that rocks the family's past,\nA reassuring heartbeat, forever meant to last.\n\nIn this dark basement, where light dare not tread,\nThe servers' gentle song, the memories are fed.\nThey nourish and sustain, through joy and sorrow's sway,\nA repository of love, in digital array.",
"_user": "Write a poem about a mass of servers humming in a dark basement, keeping a family's memories alive. At least 30 lines."
},
{
"turn": 2,
"run": 1,
"ctx_tokens_est": 60,
"new_tokens_est": 32,
"ttft": 0.147,
"gen_time": 4.242,
"gen_tps": 84.9,
"total": 4.389,
"output_tokens": 360,
"prompt_eval_count": 69,
"prompt_eval_duration_ms": 53.6,
"_response": "In a land of sun-baked stone and scrubby trees, a tortoise named Theo toiled away in his underground lair. His eyes shone bright with purpose as he tinkered with wires and circuit boards, crafting a tiny computer from scavenged parts.\n\nThe hare, Reginald, laughed at Theo's endeavor. \"What use is speed if you can't enjoy it?\" he said, sipping nectar from a nearby flower. \"My gaming rig hums like a thousand hummingbirds; it leaves your clumsy contraption in the dust!\"\n\nTheo smiled knowingly as his computer, named Terra, whirred to life with a gentle beep. Its processor creaked and groaned, but with a steady heartbeat, it churned through calculations.\n\n\"Speed may be fleeting,\" Theo said, \"but persistence is eternal.\"\n\nAnd so, while Reginald slept the day away, his rig humming in standby mode, Terra worked tirelessly, churning out masterpieces of code. It calculated pi to 10,000 digits, composed sonatas on a virtual piano, and even simulated the dance of galaxies.\n\nOne morning, as Reginald yawned awake, he beheld a stunning display: a galaxy rendered with pixel-perfect precision, swirling through space like a cosmic ballet. \"How...?\" he stammered.\n\nTheo poked his head out from his lair, grinning. \"Terra's magic, my friend. No need for haste when you have patience.\"\n\nReginald scratched his ears in wonder, and Terra continued to hum along, its processor glowing softly like a tiny star. For in this tale of tortoise and hare, it was not speed that won the day, but the steady pulse of dedication and the infinite power of persistence.",
"_user": "Write a short fable about a tortoise who built a tiny computer that was slower than the hare's gaming rig but never needed to rest."
},
{
"turn": 3,
"run": 1,
"ctx_tokens_est": 58,
"new_tokens_est": 31,
"ttft": 0.151,
"gen_time": 3.538,
"gen_tps": 82.0,
"total": 3.688,
"output_tokens": 290,
"prompt_eval_count": 68,
"prompt_eval_duration_ms": 56.9,
"_response": "In dusty attic air, where shadows play,\nA relic lies, a fading ray\nOf memory's warmth, a moment's past\nThe final print, forever to last\n\nOn yellowed paper, edges worn thin\nA face, a smile, a glimmer within\nEyes that sparkled like stars at night\nNow dim and faded, yet still alight\n\nThe subject's name, long since unknown\nLeaves only whispers, echoes thrown\nAbout a life lived, loved, and lost\nIn the silence, memories cross\n\nThe print itself, a work of art\nCrafted with care, within the heart\nOf machines that churned out scores\nOf photographs, before the digital scores\n\nBut this one stands alone, apart\nA solitary keeper of the past's heart\nA testament to moments shared\nBetween loved ones, unimpeded and unafraid\n\nThe ink is faded, colors blend\nYet in its lines, a story does tend\nTo capture love, in all its forms\nTo hold the memories that forever swarm\n\nIn attics, trunks, and boxes high\nLies hidden treasure, before our wondering eye\nThis photograph, a relic of a bygone age\nA reminder of lives lived, turned to a page\n\nWhen photographs were printed on paper's skin\nAnd love was captured in the moment within\nBefore the digital took its place\nAnd memories were lost in cyberspace.",
"_user": "Write a poem about the last photograph ever printed on paper, found in an attic a hundred years from now. At least 25 lines."
}
]
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
# Apple M5 Max / 48GB / 40 GPU cores

**Model:** llama3.1:8b (8.0B, Q4_K_M)
**Backend:** ollama
**Scenario:** creative-writing (single-shot)

| Turn | Context | Prefill | Gen | Gen tok/s | Effective tok/s | Total | Output |
|-----:|--------:|--------:|----:|----------:|----------------:|------:|-------:|
| 1 | 57 | 0.15s | 3.84s | 88.8 | **85.5** | 3.99s | 341 |
| 2 | 60 | 0.15s | 4.24s | 84.9 | **82.0** | 4.39s | 360 |
| 3 | 58 | 0.15s | 3.54s | 82.0 | **78.6** | 3.69s | 290 |

**Total prefill:** 0.4s
**Total generation:** 11.6s
**Total time:** 12.1s
**Avg generation tok/s:** 85.2
**Avg effective tok/s:** 82.1

Large diffs are not rendered by default.

Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
# Apple M5 Max / 48GB / 40 GPU cores

**Model:** llama3.1:8b (8.0B, Q4_K_M)
**Backend:** ollama
**Scenario:** doc-summary (single-shot)

| Turn | Context | Prefill | Gen | Gen tok/s | Effective tok/s | Total | Output |
|-----:|--------:|--------:|----:|----------:|----------------:|------:|-------:|
| 1 | 425 | 0.32s | 0.84s | 88.3 | **64.1** | 1.15s | 74 |
| 2 | 612 | 0.28s | 0.82s | 87.7 | **65.5** | 1.10s | 72 |
| 3 | 535 | 0.29s | 0.67s | 80.4 | **56.1** | 0.96s | 54 |
| 4 | 524 | 0.29s | 0.70s | 79.8 | **56.3** | 0.99s | 56 |
| 5 | 1,518 | 0.84s | 1.05s | 77.2 | **42.9** | 1.89s | 81 |

**Total prefill:** 2.0s
**Total generation:** 4.1s
**Total time:** 6.1s
**Avg generation tok/s:** 82.7
**Avg effective tok/s:** 55.2
161 changes: 161 additions & 0 deletions results/llama-3.1-8b-instruct/ops-agent/m5-max-48gb-40gpu_ollama.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,161 @@
{
"meta": {
"scenario": "ops-agent",
"mode": "conversation",
"label": "m5-max-48gb-40gpu ollama",
"backend": "ollama",
"model_info": {
"name": "llama3.1:8b",
"format": "gguf",
"family": "llama",
"parameter_size": "8.0B",
"quantization": "Q4_K_M"
},
"runs": 1,
"max_tokens": 500,
"cold": false,
"system": {
"os": "Darwin",
"os_version": "26.6",
"arch": "arm64",
"chip": "Apple M5 Max",
"memory_gb": 48,
"cpu_cores": 18,
"cpu_cores_performance": 6,
"cpu_cores_efficiency": 12,
"gpu_cores": 40,
"gpu_wired_limit_mb": 0
},
"timestamp": "2026-07-28T15:31:16"
},
"results": [
{
"turn": 1,
"run": 1,
"ctx_tokens_est": 575,
"new_tokens_est": 421,
"ttft": 0.42,
"gen_time": 1.87,
"gen_tps": 84.5,
"total": 2.291,
"output_tokens": 158,
"tool": "server_status",
"prompt_eval_count": 717,
"prompt_eval_duration_ms": 282.8,
"_response": "The server's been running for 12 days, which is a long time! The load average is a bit high (around 3-2.5), but it might be acceptable given the services running.\n\nLet me highlight some potential issues:\n\n* `immich-ml` is using an unusually high amount of CPU (45.2%). This could be causing performance issues or overheating.\n* `immich-postgres` and other services are running low on memory, which could lead to crashes or slowdowns if the server needs to allocate more resources.\n\nI recommend checking the service logs for `immich-ml` using `service_logs immich-ml`. Additionally, let's check the disk space usage with `disk_usage` to ensure we're not running out of storage.",
"_user": "Something feels off with the server. Can you check what's running?"
},
{
"turn": 2,
"run": 1,
"ctx_tokens_est": 979,
"new_tokens_est": 235,
"ttft": 0.281,
"gen_time": 1.907,
"gen_tps": 76.6,
"total": 2.188,
"output_tokens": 146,
"tool": "service_logs",
"prompt_eval_count": 1251,
"prompt_eval_duration_ms": 183.3,
"_response": "It looks like `immich-ml` is running a batch processing job for facial recognition and CLIP embedding. The logs suggest that it's processing a large number of images (1203 batches) and generating embeddings for each one.\n\nThe high CPU usage might be due to the complex computations involved in face detection, clustering, and embedding generation. The WARNING message about low memory condition being detected is also concerning, but the fact that the service triggered garbage collection (GC) and reclaimed 1.2GB of memory suggests it's managing its resources somewhat.\n\nHowever, I'd like to investigate further. Can you try running `docker_exec immich-ml ps` to see if there are any other processes or threads consuming CPU?",
"_user": "Immich ML is eating a lot of CPU. What's it doing?"
},
{
"turn": 3,
"run": 1,
"ctx_tokens_est": 1250,
"new_tokens_est": 89,
"ttft": 0.229,
"gen_time": 2.368,
"gen_tps": 76.0,
"total": 2.596,
"output_tokens": 180,
"tool": "docker_exec",
"prompt_eval_count": 1541,
"prompt_eval_duration_ms": 128.6,
"_response": "It looks like the `immich-ml` container is still using a significant amount of memory (8.2GB) and CPU (44.8%). The fact that it's using 68.3% of its allocated 12GB limit suggests that it might start to run low on memory soon.\n\nGiven that we're expecting this job to finish processing the vacation photos, I'd recommend keeping a close eye on the container's resource usage. If it continues to use up all available memory, it could lead to performance issues or even crashes.\n\nTo mitigate this, you could consider increasing the container's memory limit to 16GB or more, depending on your server's configuration and available resources. Alternatively, we could try to optimize the `immich-ml` configuration or investigate other solutions to reduce its resource usage.\n\nLet me know if you'd like to proceed with any of these options!",
"_user": "Ah right, we uploaded vacation photos yesterday. Is it going to be OK memory-wise?"
},
{
"turn": 4,
"run": 1,
"ctx_tokens_est": 1672,
"new_tokens_est": 214,
"ttft": 0.309,
"gen_time": 1.879,
"gen_tps": 70.3,
"total": 2.188,
"output_tokens": 132,
"tool": "backup_status",
"prompt_eval_count": 2006,
"prompt_eval_duration_ms": 207.3,
"_response": "The backup job that ran last night (2026-03-07T02:00:05) completed with warnings. It's good to know that it finished, but the warning about Vault eject failing is something we should investigate further.\n\nIt looks like the backup was able to transfer 12.4GB of data and sync 1847 files, which is a significant amount. However, the fact that Vault eject failed might indicate an issue with the backup storage or configuration.\n\nI recommend checking the backup logs for more information about the warning and potential causes. You can use `service_logs` on the backup container to see if there are any additional details.",
"_user": "Good. Check if last night's backup ran. I want to make sure the vacation photos are safe."
},
{
"turn": 5,
"run": 1,
"ctx_tokens_est": 2067,
"new_tokens_est": 239,
"ttft": 0.325,
"gen_time": 2.022,
"gen_tps": 73.2,
"total": 2.346,
"output_tokens": 148,
"tool": "disk_usage",
"prompt_eval_count": 2463,
"prompt_eval_duration_ms": 219.9,
"_response": "The disk usage is looking a bit concerning.\n\nThe root volume (`/`) has about 33% usage, which is relatively low. However, the backup vault and Time Machine volumes are both heavily used (69% and 75%, respectively). This might be due to the large amount of data being stored on these volumes.\n\nLooking at the data directory breakdown, I notice that the \"photos\" directory is taking up a significant amount of space (634GB). Given that you uploaded vacation photos yesterday, it's possible that this directory will continue to grow in size.\n\nTo free up some space, we could consider deleting or archiving older photos. We could also look into compressing or optimizing the storage of these files to reduce their footprint.",
"_user": "The vault didn't eject again. How's disk space looking on all drives?"
},
{
"turn": 6,
"run": 1,
"ctx_tokens_est": 2497,
"new_tokens_est": 249,
"ttft": 0.358,
"gen_time": 1.415,
"gen_tps": 70.7,
"total": 1.773,
"output_tokens": 100,
"tool": "service_logs",
"prompt_eval_count": 2979,
"prompt_eval_duration_ms": 252.9,
"_response": "It looks like Paperless successfully processed the school documents you scanned this morning.\n\nThe logs show that Paperless detected 3 new files in its consumption directory and processed them accordingly. It completed OCR (Optical Character Recognition) on each file, auto-classified them into relevant categories, and stored the extracted information for easy searching and organization.\n\nAll three documents were processed without any errors, so you can be confident that your school reports, certificates, and invoices are now organized and easily accessible within Paperless.",
"_user": "Time Machine is at 75%. Also did Paperless process the school documents I scanned this morning?"
},
{
"turn": 7,
"run": 1,
"ctx_tokens_est": 2740,
"new_tokens_est": 98,
"ttft": 0.212,
"gen_time": 1.526,
"gen_tps": 71.4,
"total": 1.739,
"output_tokens": 109,
"tool": "docker_exec",
"prompt_eval_count": 3202,
"prompt_eval_duration_ms": 107.5,
"_response": "The Matrix database is a bit large, but it's not unexpectedly so given the amount of activity you've been having.\n\nIt looks like most of the data is being stored in the `state_groups_state`, `event_json`, and `events` tables. If you're experiencing performance issues with your server, it might be worth considering optimizing these tables or setting up some form of database replication to reduce the load on your primary database.\n\nHowever, for now, it seems like the database is just reflecting the normal amount of activity on your Matrix instance.",
"_user": "Nice, it even picked up the electricity bill. How big is the Matrix database? It's been chatty lately."
},
{
"turn": 8,
"run": 1,
"ctx_tokens_est": 3066,
"new_tokens_est": 188,
"ttft": 0.333,
"gen_time": 2.033,
"gen_tps": 69.3,
"total": 2.366,
"output_tokens": 141,
"tool": "docker_exec",
"prompt_eval_count": 3573,
"prompt_eval_duration_ms": 221.8,
"_response": "The rooms with the most events are `#server` and `#family`, which have 12,847 and 8,934 events respectively. These two rooms seem to be contributing a significant amount of data to the Matrix database.\n\nIt might be worth taking a closer look at these rooms and seeing if there's anything that can be done to reduce the number of events being stored. This could involve adjusting the room settings or using a custom script to prune old events.\n\nThe `#test-room` also has a lot of events (4,521) which is unusual for a test room. It might be worth investigating why this is happening and whether it's contributing to the overall size of the database.",
"_user": "2.1GB for Matrix seems like a lot. Can we see which rooms take the most space?"
}
]
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
# Apple M5 Max / 48GB / 40 GPU cores

**Model:** llama3.1:8b (8.0B, Q4_K_M)
**Backend:** ollama
**Scenario:** ops-agent (conversation)

| Turn | Context | Prefill | Gen | Gen tok/s | Effective tok/s | Total | Output |
|-----:|--------:|--------:|----:|----------:|----------------:|------:|-------:|
| 1 | 575 | 0.42s | 1.87s | 84.5 | **69.0** | 2.29s | 158 |
| 2 | 979 | 0.28s | 1.91s | 76.6 | **66.7** | 2.19s | 146 |
| 3 | 1,250 | 0.23s | 2.37s | 76.0 | **69.3** | 2.60s | 180 |
| 4 | 1,672 | 0.31s | 1.88s | 70.3 | **60.3** | 2.19s | 132 |
| 5 | 2,067 | 0.33s | 2.02s | 73.2 | **63.1** | 2.35s | 148 |
| 6 | 2,497 | 0.36s | 1.42s | 70.7 | **56.4** | 1.77s | 100 |
| 7 | 2,740 | 0.21s | 1.53s | 71.4 | **62.7** | 1.74s | 109 |
| 8 | 3,066 | 0.33s | 2.03s | 69.3 | **59.6** | 2.37s | 141 |

**Total prefill:** 2.5s
**Total generation:** 15.0s
**Total time:** 17.5s
**Avg generation tok/s:** 74.0
**Avg effective tok/s:** 63.7
Loading