Skip to content

Qwen3-TTS: a fixed seed gives different audio per synthesis on one engine, and the code predictor ignores the seed #632

Description

@leehack

With the same TextToSpeechRequest and a fixed seed, back-to-back uncancelled syntheses on one loaded engine return different PCM. The first synthesis after a load is reproducible. Each later one depends on how many syntheses ran before it. The request seed also never reaches the code predictor, which samples 15 of the 16 codebooks in every frame.

Measured

Apple M4 Max, macOS 26.6.2, Flutter 3.47.1. main 8f6651f35, llamadart-native v0.4.1 (tarball sha256 41d0a529…b1487). ggml-org/Qwen3-TTS-12Hz-1.7B-Base-GGUF@ca27d74, Q4_K_M model with the Q8_0 projector. CPU (gpuLayers: 0). Text Hello from llamadart. The answer is forty two., language: 'English', maxFrames: 384. Each row is one fresh process. Hash is the first 12 hex digits of the SHA-256 of toWavBytes().

run seed temperature synthesis 1 synthesis 2 synthesis 3
A 1 0 ce4dcbc337c6 (41 frames) 2dca331835ec (32) 8a517efb689f (245)
B 2 0 ce4dcbc337c6 (41)
C 1 0.8 (default) ac7ff03787ed e7a2275b83d9 1844169b2f3d
D 2 0.8 (default) dafdcb2d4dc3
  • A: with a greedy talker, output still changes by position, and so does the frame count.
  • B equals A1: with a greedy talker, changing the seed changes nothing.
  • D differs from C1, so the seed does reach the talker.
  • C reproduces exactly in earlier CPU runs on ec6ad6f9b with v0.4.1 and on the #630 branch with v0.4.1-1.
  • On Metal (gpuLayers: 99, seed 1, temperature 0.8), all 12 processes (6 on each of those two trees) gave 37fd838eca9b, 9df98cc3f4f9, 7e790e1f8905, a80354c683c8 for syntheses 1–4. The 8 processes that ran a fifth synthesis all gave 8c9b21dd6aa5.

No cancel is involved: the post-cancel hash differences recorded in #630 occur in uncancelled syntheses too.

Cause

llamadart-native 14d4f0b (same code in v0.4.1), llama.cpp b29c606:

  • llama_dart_wrapper.cpp:211 seeds only the talker sampler (codebook 0) with request.seed.
  • llama_dart_wrapper.cpp:900-908 value-initializes mtmd_helper_gen_audio_inp and never sets input.seed, so every synthesis passes seed 0 to the code predictor. Upstream llama-tts sets it from the sampling seed (tts.cpp:123).
  • The code predictor's RNG belongs to the projector's clip_ctx (clip.cpp:177-179), so it lives as long as the loaded projector. It is reseeded only when the seed value changes (clip.cpp:4431-4434), and each frame draws one value per acoustic codebook (clip.cpp:5313-5318).
  • The first synthesis changes the seed from UINT32_MAX to 0 and reseeds. Every later synthesis also passes 0, so no reseed happens and the RNG continues from the previous synthesis. The talker KV sequence, sampler, and code2wav state are all reset per synthesis.

Causal check: a C++ harness linked against the same v0.4.1 libllamadart.dylib runs seed-1 syntheses through llama_dart_tts_* on one CPU context. Between some of them it makes one mtmd_gen_audio_process gen_code call with seed 12345, so the next synthesis's seed 0 reseeds. Both of two processes printed:

synthesis 1 2 3 (after forced reseed) 4 5 (after forced reseed)
PCM FNV-1a / frames 745b31de… / 27 3735e9da… / 33 745b31de… / 27 3735e9da… / 33 745b31de… / 27

A forced reseed restores synthesis 1 exactly and the synthesis after it matches synthesis 2, so the code-predictor RNG is the only state that carries over.

Forwarding request.seed alone would not make repeated same-seed calls match, because a repeated value does not reseed (clip.cpp:4431).

Expected vs actual

  • Expected: TextToSpeechRequest.seed is documented as "Random seed. 0xffffffff requests the runtime's random default." The same non-default seed and request should give the same PCM on every synthesis, as the first synthesis after a load already does, and the seed should drive the code predictor, as it does in llama-tts.
  • Actual: the output depends on the synthesis's position since the projector was loaded, and the seed controls only codebook 0.

Related: #603 (multimodal generate output depends on the previous request) and #322 (Qwen3-TTS hardening, which lists repeated-call evidence as open work).

Repro: tts_seed.dart, run from the repo root as dart run tts_seed.dart <model.gguf> <mmproj.gguf> <label> <seed> <temperature> <count>
import 'dart:convert';
import 'dart:io';

import 'package:crypto/crypto.dart';
import 'package:llamadart/llamadart.dart';

Future<void> main(List<String> args) async {
  final seed = int.parse(args[3]);
  final temperature = double.parse(args[4]);
  final engine = LlamaEngine(LlamaBackend());
  await engine.setLogLevel(LlamaLogLevel.none);
  await engine.loadModel(
    args[0],
    modelParams: const ModelParams(
      contextSize: 4096,
      preferredBackend: GpuBackend.cpu,
      gpuLayers: 0,
    ),
  );
  await engine.loadMultimodalProjector(args[1]);
  final tts = TextToSpeechEngine(
    engine,
    modelProfile: TextToSpeechModelProfile.qwen3Tts,
  );
  try {
    for (var i = 1; i <= int.parse(args[5]); i++) {
      final task = await tts.synthesize(
        TextToSpeechRequest(
          text: 'Hello from llamadart. The answer is forty two.',
          language: 'English',
          maxFrames: 384,
          seed: seed,
          temperature: temperature,
        ),
      );
      await for (final _ in task.events) {}
      final r = (await task.done).result!;
      stdout.writeln(jsonEncode({
        'run': args[2],
        'synthesis': i,
        'frames': r.framesGenerated,
        'sha': sha256.convert(r.toWavBytes()).toString().substring(0, 12),
      }));
    }
  } finally {
    await engine.dispose();
  }
}

Rows A–D: A 1 0 3, B 2 0 1, C 1 0.8 3, D 2 0.8 1.

Causal check: reseed.cpp

NATIVE is a llamadart-native v0.4.1 checkout, LLAMA its llama.cpp submodule (b29c606), and LIB the directory holding the v0.4.1 libllamadart.dylib (sha256 976c3cf0…):

clang++ -std=c++17 -O1 -I$NATIVE/src -I$LLAMA/include -I$LLAMA/ggml/include -I$LLAMA/tools/mtmd \
  reseed.cpp -L$LIB -lllamadart -Wl,-rpath,$LIB -o reseed
./reseed <model.gguf> <mmproj.gguf>
#include <cstdio>
#include <cstdlib>
#include <cstring>
#include <string>
#include <vector>

#include "ggml-backend.h"
#include "llama.h"
#include "llama_dart_wrapper.h"
#include "mtmd.h"

static const char *kText = "Hello from llamadart. The answer is forty two.";

static std::string synth(llama_context *lctx, mtmd_context *mctx,
                         uint32_t seed, int *frames) {
  llama_dart_tts_status st;
  llama_dart_tts *tts = llama_dart_tts_init(lctx, mctx, &st);
  if (!tts) { fprintf(stderr, "init failed %d\n", (int)st); exit(1); }
  llama_dart_tts_request req = llama_dart_tts_request_default();
  req.text = kText;
  req.text_length = strlen(kText);
  req.language = "en";
  req.max_frames = 384;
  req.seed = seed;
  if (llama_dart_tts_start(tts, &req) != LLAMA_DART_TTS_STATUS_OK) {
    fprintf(stderr, "start failed\n"); exit(1);
  }
  llama_dart_tts_progress p{};
  for (;;) {
    p.struct_size = sizeof(p);
    if (llama_dart_tts_step(tts, &p) != LLAMA_DART_TTS_STATUS_OK) {
      fprintf(stderr, "step failed\n"); exit(1);
    }
    if (p.state == LLAMA_DART_TTS_STATE_COMPLETED) break;
    if (p.state != LLAMA_DART_TTS_STATE_PROCESSING_PROMPT &&
        p.state != LLAMA_DART_TTS_STATE_GENERATING) {
      fprintf(stderr, "bad state %d\n", p.state); exit(1);
    }
  }
  *frames = p.frames_generated;
  llama_dart_tts_output_info info{};
  info.struct_size = sizeof(info);
  llama_dart_tts_get_output_info(tts, &info);
  std::vector<float> pcm((size_t)info.sample_count);
  size_t n = 0;
  llama_dart_tts_read_pcm(tts, 0, pcm.data(), pcm.size(), &n);
  uint64_t h = 1469598103934665603ULL;
  const unsigned char *b = (const unsigned char *)pcm.data();
  for (size_t i = 0; i < n * sizeof(float); i++) { h ^= b[i]; h *= 1099511628211ULL; }
  llama_dart_tts_reset(tts);
  llama_dart_tts_free(tts);
  char buf[64];
  snprintf(buf, sizeof(buf), "%016llx/%lld", (unsigned long long)h, (long long)n);
  return buf;
}

static void force_reseed(mtmd_context *mctx, uint32_t seed) {
  std::vector<float> embd(16384, 0.0f);
  mtmd_gen_inp inp = mtmd_gen_inp_default(mctx);
  inp.type = MTMD_GEN_PROCESS_TYPE_GEN_CODE;
  inp.code0 = 0;
  inp.embd = embd.data();
  inp.seed = seed;
  mtmd_gen_out out{};
  if (mtmd_gen_audio_process(mctx, &inp, &out) != 0) {
    fprintf(stderr, "force_reseed failed\n"); exit(1);
  }
}

int main(int argc, char **argv) {
  if (argc < 3) return 2;
  llama_log_set([](ggml_log_level, const char *, void *) {}, nullptr);
  ggml_backend_load_all();
  llama_backend_init();
  ggml_backend_dev_t cpu = ggml_backend_dev_by_type(GGML_BACKEND_DEVICE_TYPE_CPU);
  ggml_backend_dev_t devs[2] = {cpu, nullptr};
  llama_model_params mp = llama_model_default_params();
  mp.n_gpu_layers = 0;
  mp.devices = devs;
  llama_model *model = llama_model_load_from_file(argv[1], mp);
  if (!model) return 3;
  llama_context_params cp = llama_context_default_params();
  cp.n_ctx = 4096;
  cp.op_offload = false;
  llama_context *lctx = llama_init_from_model(model, cp);
  if (!lctx) return 4;
  mtmd_context_params mcp = mtmd_context_params_default();
  mcp.use_gpu = false;
  mcp.print_timings = false;
  mtmd_context *mctx = mtmd_init_from_file(argv[2], model, mcp);
  if (!mctx) return 5;
  llama_set_embeddings(lctx, true);
  int f = 0;
  std::string s;
  s = synth(lctx, mctx, 1, &f); printf("s1 plain           %s frames=%d\n", s.c_str(), f); fflush(stdout);
  s = synth(lctx, mctx, 1, &f); printf("s2 plain           %s frames=%d\n", s.c_str(), f); fflush(stdout);
  force_reseed(mctx, 12345);
  s = synth(lctx, mctx, 1, &f); printf("s3 after reseed    %s frames=%d\n", s.c_str(), f); fflush(stdout);
  s = synth(lctx, mctx, 1, &f); printf("s4 plain           %s frames=%d\n", s.c_str(), f); fflush(stdout);
  force_reseed(mctx, 12345);
  s = synth(lctx, mctx, 1, &f); printf("s5 after reseed    %s frames=%d\n", s.c_str(), f); fflush(stdout);
  mtmd_free(mctx);
  llama_free(lctx);
  llama_model_free(model);
  return 0;
}

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    blocked-runtimeRequires upstream or native runtime support before Dart package work can complete.bugSomething isn't workingpriority:P2Planned next: useful unblocked work or validation after P1 items

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions