https://github.com/ollama/ollama/issues/7518 trying to figure out why I get only 1 generation from ollama. I suspect this is due to constraints of its api, but noting here for future discussion perhaps related to this same issue with llama.cpp server? https://github.com/cosmicoptima/loom/pull/25
ollama/ollama#7518
trying to figure out why I get only 1 generation from ollama.
I suspect this is due to constraints of its api, but noting here for future discussion
perhaps related to this same issue with llama.cpp server? #25