Skip to content

Server silently serves the resident model for unresolvable model tags, echoing the requested name #716

Description

@Platano78

Summary

flm serve answers requests for model tags it cannot resolve by running whatever model is
currently loaded
, and echoes the requested tag back in the response model field. Nothing in
the response or the server log indicates the substitution. A benchmark script can therefore collect
timings for a model that never loaded.

Reproduce (10 seconds)

flm serve qwen3:1.7b --host 127.0.0.1 --port 8093

curl -s http://127.0.0.1:8093/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"model":"totally-made-up:99b","messages":[{"role":"user","content":"Name three colors."}],"max_tokens":40}'

Actual — a fully formed completion for a tag that exists nowhere:

{"model":"totally-made-up:99b",
 "choices":[{"message":{"role":"assistant","content":"Here are three colors:\n\n1. Red\n2. Blue\n3. Green"}}],
 "usage":{"decoding_speed_tps":43.27, ...}}

Expected: an error naming the unresolvable tag.

Same on /v1/audio/transcriptions: one loaded Whisper answered to whisper-v3:turbo,
whisper-v3-turbo-FLM and whisper-1 alike, each echoed back verbatim.

Why it matters

The model field is the only signal a client has about which model ran, and it is an echo of the
request rather than a statement of fact. The failure is easiest to hit exactly when a tag should
fail — mid-setup, before model_list.json is correct — which is also when people run their first
benchmark. The result looks completely normal.

Note this is distinct from a registered-but-unloadable tag, which does error correctly.

Suggested fix

Return an error for an unresolvable tag instead of falling through to the resident model. Failing
that, populate the response model field with the model that actually served the request, so the
output is self-describing.

Environment: FLM v1.0.1 / v1.0.2 / v1.0.4, Ryzen AI Max+ 395, NPU fw 1.1.2.65, in-tree
amdxdna 0.7, kernel 7.0.0-31-generic, Ubuntu 26.04.

Activity

  1. Atomic-Germ commented on Sep 9, 2026

    @Atomic-Germ

    Ah yes; I went around in circles for so So long when I first started making custom models, because of that.

  2. noamsto commented on Sep 28, 2026

    @noamsto

    We see this on v1.0.6 (Strix Halo, Linux), and one detail matters for the fix. The request is not served by the resident model. flm serve evicts the resident model, loads llama3.2:1b in its place, and answers under the requested name.

    With flm serve gemma4-it:e4b:

    $ curl -s :52625/api/ps | jq -c '.models[0].name'
    "gemma4-it:e4b"
    $ curl -s :52625/v1/chat/completions -H 'Content-Type: application/json' \
        -d '{"model":"oflm-test-no-such-model:0b","messages":[{"role":"user","content":"Say ok."}],"max_tokens":8}'
    {"model":"oflm-test-no-such-model:0b","choices":[{"message":{"content":"Ok"}}],"usage":{"load_duration":2.499, ...}}   HTTP 200
    $ curl -s :52625/api/ps | jq -c '.models[0].name'
    "llama3.2:1b"
    

    The server log for that request shows Model tag 'oflm-test-no-such-model:0b' is not supported, followed by Loading model: .../Llama-3.2-1B-NPU2. The #716 repro served qwen3:1.7b, so the answer there most likely came from llama3.2:1b, not qwen3.

    The path:

    • ensure_model_loaded resets the engine before it resolves the tag (rest_handler.cpp:382-388).
    • get_auto_model falls back to Llama3 for an unsupported tag (all_models.hpp:92-95). It does the same for a known non-chat family such as embed-gemma (:174-178).
    • The handler echoes the requested name.

    We saw the same substitution for "model": "" and for the internal "model-faker" sentinel, in both streaming and non-streaming mode. From the source, without measuring it: current_model_tag is then set to llama3.2:1b (:418), so the next request with the same bad tag fails the comparison at :382 and reloads again.

    /v1/embeddings has the embedding version of the same problem. It never checks model against the loaded embedding model, and it echoes the requested name next to the loaded model's real vectors (:882, :914).

    OpenFlowLM-Next fixed this as its SERVER-MODEL-IDENTITY requirement. We carry a port of that fix on v1.0.6:
    https://github.com/noamsto/nix-amd-ai/blob/46d7a73c1044435cec0ff08410e1a3bc87cbc544/pkgs/fastflowlm/patches/model-identity.patch

    It resolves the tag before anything is unloaded, and handles each case like this:

    • An unknown tag, "", "model-faker", or a non-chat tag on a chat route gets 400 model_not_found, and the served model stays loaded.
    • A known model that fails to load gets 500 model_load_failed.
    • /v1/embeddings refuses any tag but the loaded one.

    The logs, including a before/after /api/ps check that the model is not evicted, are at https://github.com/noamsto/nix-amd-ai/blob/46d7a73c1044435cec0ff08410e1a3bc87cbc544/bench-logs/oflm-api-conformance-2026-09-28-after-173/README.md

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions