Qwen, DeepSeek, Kimi and GLM: what to self-host?
Compare Qwen3.8, DeepSeek V4.1, Kimi K3 and GLM-5.3. Understand weights, hardware needs and model selection for private deployment with Syntalith.
Syntalith
Qwen on an employee’s computer and Kimi on a GPU cluster can both use open weights. Their budgets, operating requirements and suitable tasks are very different. A useful model comparison therefore starts with where the system will run and what work it must complete.
This shortlist uses publisher material checked on 28 September 2026. These are candidates for evaluation. We have not run a shared benchmark of all these models on client data.
What the current families offer
| Family and version | What the primary source establishes | Deployment implication |
|---|---|---|
| Qwen3.8-27B | Downloadable weights for a smaller Qwen3.8 variant | A candidate for a controlled single-GPU installation with quantization |
| Qwen3.8-Flash-Next | A 125B backbone, 6B active parameters, plus n-gram embeddings and MTP | The active count does not describe the files or total memory requirement |
| DeepSeek-V4.1-Flash | Multimodal MoE with a 552B backbone, released under MIT | Assess the full model and whether the server supports its architecture |
| Kimi K3 | 2.8T parameters, 104B active, its own Kimi K3 licence | Infrastructure must accommodate very large weights |
| GLM-5.3 and GLM-5.3-Flash | The repository lists 744B-A40B and 320B-A18B variants and local serving options | Check the version, precision and licence of the actual download |
Qwen has also published its Qwen3.8-Omni report on text, images, audio and video. A research report and released tools do not establish that every described model is downloadable. A deployment quote needs actual weights and a working serving engine.
Keep smaller candidates in the test
Gemma 4 12B, Mistral Small 4 and gpt-oss cover different sizes and architectures. A name such as Small does not specify consumer hardware requirements: Mistral Small 4 has 119B total parameters and 6.5B active.
For Polish documents, we also consider Bielik and PLLuM. Our comparison with Qwen identifies particular versions and licensing constraints. A focus on Polish is a reason to include a model in the trial. Your own questions establish how well it handles your terminology.
Open source, open weights and private data
Access to weights lets you choose the inference environment. A licence determines permitted use, modification and distribution. The serving code can have a different licence from the model. We record the model identifier, revision and terms associated with the selected files.
Privacy also depends on the application. A local model connected to cloud OCR still sends a document to an external service. Embeddings, backups and diagnostics need separate checks. Our private LLM deployment service considers the complete path.
How Syntalith selects a model
For a help desk, we measure routing accuracy and required corrections. A legal department needs accurate clause references and abstention when a document is missing. A coding task needs a working change and tests. One general leaderboard cannot answer all these questions.
We start with a short list that fits the budget. Each candidate receives the same material and acceptance criteria. The comparison covers output quality, latency under the intended load and cost per successfully completed case. A model needing fewer corrections may justify higher compute costs.
Send us your use case, data requirements and hardware specifications. Syntalith can scope the comparison, then install, host and adapt the selected model for a business or personal setup.
Evaluate private AI for your organization
We help businesses and individuals select hardware, deploy a model and test it on their own tasks. Start with a computer you already own or ask us before buying one.
Private LLMs and fine-tuning