When not to deploy Qwen locally: five results that should stop a project
A 121-minute non-result, three hours of compaction, 250k without better quality, missed validation and a broken frontend. This is an evidence-based no-build decision.
Syntalith
One run in the local experiment lasted 121 minutes 38 seconds and delivered no working change. Another profile needed more than three hours and two user interventions, then still broke an existing contract. Those are sufficient reasons to reject a setup even when the model fits on the card and looks good in a short demo.
This is not a generic list of local-LLM disadvantages. Each case below comes from the after-hours home-PC experiment with an RTX 3090 and maps to a decision a buyer should make.
Result 1: 121 minutes, 4.28 million input tokens and no completion
Claude Code 2.1.235 was connected to Qwen3.8-27B through a local Anthropic-compatible endpoint. The llama.cpp profile had a 120k context and medium effort. The agent received normal Bash, Edit, Read, Write, Glob and Grep tools in a clean writable worktree.
There was no artificial wall-time cap. After 7,298.5 seconds, the client submitted 131,265 tokens to a 131,072-token server. The server returned HTTP 400. By then the harness had made 66 tool calls and cumulatively resent 4,283,936 input tokens. Its 16-test file did not compile, documentation was untouched and no final report existed.
This was not evidence against Qwen as a model, nor against Claude Code with Claude models. It rejected this unsupported combination of harness, endpoint and 120k context as the experiment's reference profile.
Deployment decision: when a representative task does not finish naturally, buying more hardware is not the first fix. Inspect context management, repeated prefill, protocol and harness behaviour. The right move may be a different client or a managed API.
Result 2: a small context turned one task into a three-hour session
Qwen Code medium on the corrected 60k profile received an ordinary handoff-error task. Its first turn lasted 152.99 minutes and entered this cycle:
generate the complete test file
-> hit output or context limit
-> automatic summary
-> lose read-before-write state
-> reread
-> generate the complete file again
Two natural user corrections eventually finished the session. Combined time was about 192.1 minutes. The diff became disciplined only after an explicit request to restore out-of-scope files. It still rejected an empty REST body that had been valid existing behaviour. Independent rubric score: 78/100.
The same kind of task on a better 120k/medium profile finished much sooner. A larger context does not guarantee quality, but an undersized working window can destroy agent throughput.
Deployment decision: if the session repeatedly compacts before the first material patch, the profile is not ready. Record compactions, rereads and human interventions alongside token speed.
Result 3: maximum context increased cost without improving quality
The 250k profile avoided compaction and retained about 119k of live context. It completed in 30 minutes 19 seconds and peaked at 23,179 MiB of VRAM. It sounded like the obvious winner, but it also broke empty-body handling and scored 80/100. The 120k/medium profile preserved compatibility and remained the better decision.
In a separate recall check, a 230,085-token prompt passed all 12 facts but took 563.4 seconds. A 50,059-token version also passed 12/12 and took 59.3 seconds. The maximum-window run was almost ten times slower in this cell.
Deployment decision: if the larger window does not improve process acceptance, shorten the packet, add retrieval or split the workflow. Do not pay the 250k latency tax merely because the checkpoint accepts it.
Result 4: public tests passed and hidden validation did not
For CSV import, Qwen Code medium finished in 93.14 seconds, passed 12/12 public checks and 4/5 hidden checks. It failed to reject a blank name and role. Codex and OpenCode driving the same local model passed every check for this task.
The defect is small inside a fixture. In account provisioning it could create invalid records, manual cleanup and a difficult audit trail.
Deployment decision: without edge cases, there is no production evidence. If a costly defect cannot be caught automatically or reviewed cheaply, use a narrower system, stronger model or deterministic component.
Result 5: impressive generation failed browser acceptance
Qwen Code generated an ATS frontend from an empty folder in 26:03. It had search, sorting, stages, a detail drawer and 10,000 synthetic applicants. Its virtual list had no bounded viewport, so the page rendered 8,652 matching rows and 251,327 DOM elements. The initial run also ended on a false loop-detector signal before full browser QA.
Repairing six defects and later fixing two visual issues brought total model time to 48:25. Only then did independent acceptance confirm bounded rows, no page overflow and correct desktop/mobile geometry.
Deployment decision: a demo without outcome verification remains a prototype. UI work needs behaviour, performance, responsive checks and image review. Document systems need equivalent checks for citations, evidence-free refusal and permissions.
Project stop table
| Pilot signal | Action before more investment |
|---|---|
| task does not finish or overflows context | change harness, profile or model; rerun the same fixture |
| compaction and rereads dominate | enlarge the useful window or reduce tools and source material |
| longer context does not improve acceptance | use retrieval, chunking or a shorter workflow |
| an edge-case error carries high cost | expand the suite or use a deterministic component |
| output looks good but lacks an acceptance test | do not admit users |
| no owner exists for upgrades and incidents | buy managed operations or stop |
What a buyer receives before committing to a build
Syntalith can run a short rejection-oriented evaluation: choose one process, build 20–50 cases, test a local checkpoint and a sensible API baseline, and measure time, cost, corrections and risk. The result may favour local Qwen, a smaller model, deterministic code, an API or no project.
The free process scan selects the case. An AI process audit delivers the comparison, acceptance rules, architecture and fixed quote. The product is a defensible decision before infrastructure spend. A model chosen in advance receives no preferential treatment.
Free process scan
Start with a free process scan.
- A 30-minute call with the engineer who would lead the work.
- A review of the processes that cost you the most time and money.
- A written summary of what to automate first and the likely cost range.
The scan chooses one process to assess, and within 2 business days you receive a recommendation, including when a simpler route is the better fit.
€0
30 minutes · written takeaway within 2 business days
Times are shown in your own time zone. We work with clients across time zones.
Describe the process in the form