#testing
-
mistral.rs 0.8’s generate_structured hangs on GGUF, so I parse the JSON myself
Grammar-constrained decoding produced about zero tokens in three minutes while plain chat ran at 25 tokens per second. Here’s what LocalGPT Verse ships instead: plain instructed JSON, a 30-line parser, serde defaults, clamps, and a fallback.