With the same TextToSpeechRequest and a fixed seed, back-to-back uncancelled syntheses on one loaded engine return different PCM. The first synthesis after a load is reproducible. Each later one depends on how many syntheses ran before it. The request seed also never reaches the code predictor, which samples 15 of the 16 codebooks in every frame.
Measured
Apple M4 Max, macOS 26.6.2, Flutter 3.47.1. main 8f6651f35, llamadart-native v0.4.1 (tarball sha256 41d0a529…b1487). ggml-org/Qwen3-TTS-12Hz-1.7B-Base-GGUF@ca27d74, Q4_K_M model with the Q8_0 projector. CPU (gpuLayers: 0). Text Hello from llamadart. The answer is forty two., language: 'English', maxFrames: 384. Each row is one fresh process. Hash is the first 12 hex digits of the SHA-256 of toWavBytes().
| run |
seed |
temperature |
synthesis 1 |
synthesis 2 |
synthesis 3 |
| A |
1 |
0 |
ce4dcbc337c6 (41 frames) |
2dca331835ec (32) |
8a517efb689f (245) |
| B |
2 |
0 |
ce4dcbc337c6 (41) |
|
|
| C |
1 |
0.8 (default) |
ac7ff03787ed |
e7a2275b83d9 |
1844169b2f3d |
| D |
2 |
0.8 (default) |
dafdcb2d4dc3 |
|
|
- A: with a greedy talker, output still changes by position, and so does the frame count.
- B equals A1: with a greedy talker, changing the seed changes nothing.
- D differs from C1, so the seed does reach the talker.
- C reproduces exactly in earlier CPU runs on
ec6ad6f9b with v0.4.1 and on the #630 branch with v0.4.1-1.
- On Metal (
gpuLayers: 99, seed 1, temperature 0.8), all 12 processes (6 on each of those two trees) gave 37fd838eca9b, 9df98cc3f4f9, 7e790e1f8905, a80354c683c8 for syntheses 1–4. The 8 processes that ran a fifth synthesis all gave 8c9b21dd6aa5.
No cancel is involved: the post-cancel hash differences recorded in #630 occur in uncancelled syntheses too.
Cause
llamadart-native 14d4f0b (same code in v0.4.1), llama.cpp b29c606:
llama_dart_wrapper.cpp:211 seeds only the talker sampler (codebook 0) with request.seed.
llama_dart_wrapper.cpp:900-908 value-initializes mtmd_helper_gen_audio_inp and never sets input.seed, so every synthesis passes seed 0 to the code predictor. Upstream llama-tts sets it from the sampling seed (tts.cpp:123).
- The code predictor's RNG belongs to the projector's
clip_ctx (clip.cpp:177-179), so it lives as long as the loaded projector. It is reseeded only when the seed value changes (clip.cpp:4431-4434), and each frame draws one value per acoustic codebook (clip.cpp:5313-5318).
- The first synthesis changes the seed from
UINT32_MAX to 0 and reseeds. Every later synthesis also passes 0, so no reseed happens and the RNG continues from the previous synthesis. The talker KV sequence, sampler, and code2wav state are all reset per synthesis.
Causal check: a C++ harness linked against the same v0.4.1 libllamadart.dylib runs seed-1 syntheses through llama_dart_tts_* on one CPU context. Between some of them it makes one mtmd_gen_audio_process gen_code call with seed 12345, so the next synthesis's seed 0 reseeds. Both of two processes printed:
| synthesis |
1 |
2 |
3 (after forced reseed) |
4 |
5 (after forced reseed) |
| PCM FNV-1a / frames |
745b31de… / 27 |
3735e9da… / 33 |
745b31de… / 27 |
3735e9da… / 33 |
745b31de… / 27 |
A forced reseed restores synthesis 1 exactly and the synthesis after it matches synthesis 2, so the code-predictor RNG is the only state that carries over.
Forwarding request.seed alone would not make repeated same-seed calls match, because a repeated value does not reseed (clip.cpp:4431).
Expected vs actual
- Expected:
TextToSpeechRequest.seed is documented as "Random seed. 0xffffffff requests the runtime's random default." The same non-default seed and request should give the same PCM on every synthesis, as the first synthesis after a load already does, and the seed should drive the code predictor, as it does in llama-tts.
- Actual: the output depends on the synthesis's position since the projector was loaded, and the seed controls only codebook 0.
Related: #603 (multimodal generate output depends on the previous request) and #322 (Qwen3-TTS hardening, which lists repeated-call evidence as open work).
Repro: tts_seed.dart, run from the repo root as dart run tts_seed.dart <model.gguf> <mmproj.gguf> <label> <seed> <temperature> <count>
import 'dart:convert';
import 'dart:io';
import 'package:crypto/crypto.dart';
import 'package:llamadart/llamadart.dart';
Future<void> main(List<String> args) async {
final seed = int.parse(args[3]);
final temperature = double.parse(args[4]);
final engine = LlamaEngine(LlamaBackend());
await engine.setLogLevel(LlamaLogLevel.none);
await engine.loadModel(
args[0],
modelParams: const ModelParams(
contextSize: 4096,
preferredBackend: GpuBackend.cpu,
gpuLayers: 0,
),
);
await engine.loadMultimodalProjector(args[1]);
final tts = TextToSpeechEngine(
engine,
modelProfile: TextToSpeechModelProfile.qwen3Tts,
);
try {
for (var i = 1; i <= int.parse(args[5]); i++) {
final task = await tts.synthesize(
TextToSpeechRequest(
text: 'Hello from llamadart. The answer is forty two.',
language: 'English',
maxFrames: 384,
seed: seed,
temperature: temperature,
),
);
await for (final _ in task.events) {}
final r = (await task.done).result!;
stdout.writeln(jsonEncode({
'run': args[2],
'synthesis': i,
'frames': r.framesGenerated,
'sha': sha256.convert(r.toWavBytes()).toString().substring(0, 12),
}));
}
} finally {
await engine.dispose();
}
}
Rows A–D: A 1 0 3, B 2 0 1, C 1 0.8 3, D 2 0.8 1.
Causal check: reseed.cpp
NATIVE is a llamadart-native v0.4.1 checkout, LLAMA its llama.cpp submodule (b29c606), and LIB the directory holding the v0.4.1 libllamadart.dylib (sha256 976c3cf0…):
clang++ -std=c++17 -O1 -I$NATIVE/src -I$LLAMA/include -I$LLAMA/ggml/include -I$LLAMA/tools/mtmd \
reseed.cpp -L$LIB -lllamadart -Wl,-rpath,$LIB -o reseed
./reseed <model.gguf> <mmproj.gguf>
#include <cstdio>
#include <cstdlib>
#include <cstring>
#include <string>
#include <vector>
#include "ggml-backend.h"
#include "llama.h"
#include "llama_dart_wrapper.h"
#include "mtmd.h"
static const char *kText = "Hello from llamadart. The answer is forty two.";
static std::string synth(llama_context *lctx, mtmd_context *mctx,
uint32_t seed, int *frames) {
llama_dart_tts_status st;
llama_dart_tts *tts = llama_dart_tts_init(lctx, mctx, &st);
if (!tts) { fprintf(stderr, "init failed %d\n", (int)st); exit(1); }
llama_dart_tts_request req = llama_dart_tts_request_default();
req.text = kText;
req.text_length = strlen(kText);
req.language = "en";
req.max_frames = 384;
req.seed = seed;
if (llama_dart_tts_start(tts, &req) != LLAMA_DART_TTS_STATUS_OK) {
fprintf(stderr, "start failed\n"); exit(1);
}
llama_dart_tts_progress p{};
for (;;) {
p.struct_size = sizeof(p);
if (llama_dart_tts_step(tts, &p) != LLAMA_DART_TTS_STATUS_OK) {
fprintf(stderr, "step failed\n"); exit(1);
}
if (p.state == LLAMA_DART_TTS_STATE_COMPLETED) break;
if (p.state != LLAMA_DART_TTS_STATE_PROCESSING_PROMPT &&
p.state != LLAMA_DART_TTS_STATE_GENERATING) {
fprintf(stderr, "bad state %d\n", p.state); exit(1);
}
}
*frames = p.frames_generated;
llama_dart_tts_output_info info{};
info.struct_size = sizeof(info);
llama_dart_tts_get_output_info(tts, &info);
std::vector<float> pcm((size_t)info.sample_count);
size_t n = 0;
llama_dart_tts_read_pcm(tts, 0, pcm.data(), pcm.size(), &n);
uint64_t h = 1469598103934665603ULL;
const unsigned char *b = (const unsigned char *)pcm.data();
for (size_t i = 0; i < n * sizeof(float); i++) { h ^= b[i]; h *= 1099511628211ULL; }
llama_dart_tts_reset(tts);
llama_dart_tts_free(tts);
char buf[64];
snprintf(buf, sizeof(buf), "%016llx/%lld", (unsigned long long)h, (long long)n);
return buf;
}
static void force_reseed(mtmd_context *mctx, uint32_t seed) {
std::vector<float> embd(16384, 0.0f);
mtmd_gen_inp inp = mtmd_gen_inp_default(mctx);
inp.type = MTMD_GEN_PROCESS_TYPE_GEN_CODE;
inp.code0 = 0;
inp.embd = embd.data();
inp.seed = seed;
mtmd_gen_out out{};
if (mtmd_gen_audio_process(mctx, &inp, &out) != 0) {
fprintf(stderr, "force_reseed failed\n"); exit(1);
}
}
int main(int argc, char **argv) {
if (argc < 3) return 2;
llama_log_set([](ggml_log_level, const char *, void *) {}, nullptr);
ggml_backend_load_all();
llama_backend_init();
ggml_backend_dev_t cpu = ggml_backend_dev_by_type(GGML_BACKEND_DEVICE_TYPE_CPU);
ggml_backend_dev_t devs[2] = {cpu, nullptr};
llama_model_params mp = llama_model_default_params();
mp.n_gpu_layers = 0;
mp.devices = devs;
llama_model *model = llama_model_load_from_file(argv[1], mp);
if (!model) return 3;
llama_context_params cp = llama_context_default_params();
cp.n_ctx = 4096;
cp.op_offload = false;
llama_context *lctx = llama_init_from_model(model, cp);
if (!lctx) return 4;
mtmd_context_params mcp = mtmd_context_params_default();
mcp.use_gpu = false;
mcp.print_timings = false;
mtmd_context *mctx = mtmd_init_from_file(argv[2], model, mcp);
if (!mctx) return 5;
llama_set_embeddings(lctx, true);
int f = 0;
std::string s;
s = synth(lctx, mctx, 1, &f); printf("s1 plain %s frames=%d\n", s.c_str(), f); fflush(stdout);
s = synth(lctx, mctx, 1, &f); printf("s2 plain %s frames=%d\n", s.c_str(), f); fflush(stdout);
force_reseed(mctx, 12345);
s = synth(lctx, mctx, 1, &f); printf("s3 after reseed %s frames=%d\n", s.c_str(), f); fflush(stdout);
s = synth(lctx, mctx, 1, &f); printf("s4 plain %s frames=%d\n", s.c_str(), f); fflush(stdout);
force_reseed(mctx, 12345);
s = synth(lctx, mctx, 1, &f); printf("s5 after reseed %s frames=%d\n", s.c_str(), f); fflush(stdout);
mtmd_free(mctx);
llama_free(lctx);
llama_model_free(model);
return 0;
}
With the same
TextToSpeechRequestand a fixedseed, back-to-back uncancelled syntheses on one loaded engine return different PCM. The first synthesis after a load is reproducible. Each later one depends on how many syntheses ran before it. The request seed also never reaches the code predictor, which samples 15 of the 16 codebooks in every frame.Measured
Apple M4 Max, macOS 26.6.2, Flutter 3.47.1. main
8f6651f35, llamadart-nativev0.4.1(tarball sha25641d0a529…b1487). ggml-org/Qwen3-TTS-12Hz-1.7B-Base-GGUF@ca27d74, Q4_K_M model with the Q8_0 projector. CPU (gpuLayers: 0). TextHello from llamadart. The answer is forty two.,language: 'English',maxFrames: 384. Each row is one fresh process. Hash is the first 12 hex digits of the SHA-256 oftoWavBytes().ce4dcbc337c6(41 frames)2dca331835ec(32)8a517efb689f(245)ce4dcbc337c6(41)ac7ff03787ede7a2275b83d91844169b2f3ddafdcb2d4dc3ec6ad6f9bwithv0.4.1and on the #630 branch withv0.4.1-1.gpuLayers: 99, seed 1, temperature 0.8), all 12 processes (6 on each of those two trees) gave37fd838eca9b,9df98cc3f4f9,7e790e1f8905,a80354c683c8for syntheses 1–4. The 8 processes that ran a fifth synthesis all gave8c9b21dd6aa5.No cancel is involved: the post-cancel hash differences recorded in #630 occur in uncancelled syntheses too.
Cause
llamadart-native
14d4f0b(same code inv0.4.1), llama.cppb29c606:llama_dart_wrapper.cpp:211seeds only the talker sampler (codebook 0) withrequest.seed.llama_dart_wrapper.cpp:900-908value-initializesmtmd_helper_gen_audio_inpand never setsinput.seed, so every synthesis passes seed0to the code predictor. Upstreamllama-ttssets it from the sampling seed (tts.cpp:123).clip_ctx(clip.cpp:177-179), so it lives as long as the loaded projector. It is reseeded only when the seed value changes (clip.cpp:4431-4434), and each frame draws one value per acoustic codebook (clip.cpp:5313-5318).UINT32_MAXto0and reseeds. Every later synthesis also passes0, so no reseed happens and the RNG continues from the previous synthesis. The talker KV sequence, sampler, and code2wav state are all reset per synthesis.Causal check: a C++ harness linked against the same
v0.4.1libllamadart.dylibruns seed-1 syntheses throughllama_dart_tts_*on one CPU context. Between some of them it makes onemtmd_gen_audio_processgen_code call with seed 12345, so the next synthesis's seed0reseeds. Both of two processes printed:745b31de…/ 273735e9da…/ 33745b31de…/ 273735e9da…/ 33745b31de…/ 27A forced reseed restores synthesis 1 exactly and the synthesis after it matches synthesis 2, so the code-predictor RNG is the only state that carries over.
Forwarding
request.seedalone would not make repeated same-seed calls match, because a repeated value does not reseed (clip.cpp:4431).Expected vs actual
TextToSpeechRequest.seedis documented as "Random seed.0xffffffffrequests the runtime's random default." The same non-default seed and request should give the same PCM on every synthesis, as the first synthesis after a load already does, and the seed should drive the code predictor, as it does inllama-tts.Related: #603 (multimodal generate output depends on the previous request) and #322 (Qwen3-TTS hardening, which lists repeated-call evidence as open work).
Repro:
tts_seed.dart, run from the repo root asdart run tts_seed.dart <model.gguf> <mmproj.gguf> <label> <seed> <temperature> <count>Rows A–D:
A 1 0 3,B 2 0 1,C 1 0.8 3,D 2 0.8 1.Causal check:
reseed.cppNATIVEis a llamadart-nativev0.4.1checkout,LLAMAits llama.cpp submodule (b29c606), andLIBthe directory holding thev0.4.1libllamadart.dylib(sha256976c3cf0…):