Polish company RAG: two studies, two different winners
Bielik led one Polish RAG study and PLLuM led another. Here are the numbers, why the order reversed, and how to build an auditable company brain.
Syntalith
A “company brain” is not a model with a document folder poured into it. It is a system that must retrieve the right source, apply permissions, use the effective document version and expose evidence. Only then should the team choose Qwen, Bielik, PLLuM or an API model as the generator.
Two 2026 papers show why that component cannot be selected by brand. Bielik led a RAG study over technical documentation. PLLuM led another study over Polish Wikipedia. Both results are real and cover different checkpoints and pipelines.
Study one: real technical documentation
Evaluation of Two Leading Polish Language Models in a Real-world RAG Scenario used real technical documentation for a low-code platform, split into about 1,200 chunks. The evaluation contained 146 reference questions selected as likely user questions, with employees helping to build the reference answers.
The best vector retrieval setup, using OrlikB/KartonBERT-USE-base-v1, produced:
| Retrieved passages | Accuracy | Recall | F1 | NDCG |
|---|---|---|---|---|
| 5 | 0.891 | 0.744 | 0.485 | 0.682 |
| 7 | 0.899 | 0.796 | 0.424 | 0.701 |
The generator received five documents. Mean scores on a 1–5 scale were:
| Checkpoint | Mean score |
|---|---|
| Bielik-11B-v2.3-Instruct | 4.521 |
| PLLuM-12B-nc-chat | 4.025 |
The study also exposed judge bias. When PLLuM's answer appeared first, Bielik won 81.5% of comparisons. After answer order was reversed, Bielik's share fell to 54.3%. Pairwise evaluation without order randomisation can therefore manufacture a large apparent margin.
Study two: Polish Wikipedia
Evaluating Cost-Efficiency of LLMs in a RAG Setup on Polish Wikipedia used 1,000 PolQA questions. Its pipeline retrieved candidates from Polish Wikipedia, reranked them and sent the selected passages to the generator. GPT-4o provided pairwise judgments.
The Bradley-Terry ratings were:
| Checkpoint | Rating |
|---|---|
| PLLuM-12B-nc-chat FP16 | 1.273 ± 0.030 |
| Bielik-11B-v2.6-Instruct FP16 | 1.061 ± 0.014 |
PLLuM led this study. That does not invalidate the first paper. The Bielik checkpoint, corpus, questions, retriever, reranker and judging procedure all changed. “The best Polish RAG model” is not a useful claim without those details.
Where Qwen3.8-27B fits
The official Qwen3.8-27B model card describes a dense 27B model with a native 262,144-token window. In our home experiment, Qwen recalled every checked fact from prompts containing 50,059, 115,074 and 230,085 tokens. The longest run took 563.4 seconds and peaked at 23,623 MiB of VRAM.
That result places Qwen on the shortlist for long-document testing. It does not settle the generator choice for Polish RAG. Our Qwen run only checked whether the model could retrieve three facts from a long input; the RAG scores above belong only to the checkpoints used in those papers.
A knowledge system has at least two results
The first result is retrieval: did the right passage rank high enough? The second is generation: did the model preserve the source meaning, citation, version and justified refusal?
A single aggregate score hides the failure location. When an answer is wrong, the team should be able to determine whether:
- parsing lost a table or heading;
- chunking separated a rule from its exception;
- retrieval missed the effective document;
- an access filter removed an allowed source or admitted a forbidden one;
- generation ignored the supplied passage;
- the index retained an obsolete version.
Evidence retained with every answer
An answer without a source trace is difficult to accept and difficult to repair. An operational system should retain document identity and version, ranked passages, user identity and access-filter result, exact checkpoint and profile, final answer and reviewer decision.
That record allows the case to be repeated after a parser, embedding or model change. Without it, an upgrade becomes an impression contest.
A pre-deployment evaluation
Build the set from questions employees actually ask. The process owner identifies the governing source, effective version, acceptable answer and refusal cases. Access-control cases belong in the same set rather than being added after model selection.
Compare exact profiles on the same index. If retrieval changes, score retrieval separately from the generator. The report should show every rejected answer and its reason. An average alone is insufficient.
What Syntalith can build and transfer
We can start with a source audit and question set, then implement parsing, indexing, role filters, citations, a review panel and an upgrade gate. We compare local models with a sensible API alternative. If Qwen wins on the buyer's task, we deploy Qwen. If Bielik, PLLuM or a simpler non-LLM system wins, the recommendation should say so.
See our AI process audit and custom AI applications, or begin with a free process scan. Delivery includes implementation, acceptance evidence, documentation and training for the people who will own the system.
Free process scan
Start with a free process scan.
- A 30-minute call with the engineer who would lead the work.
- A review of the processes that cost you the most time and money.
- A written summary of what to automate first and the likely cost range.
The scan chooses one process to assess, and within 2 business days you receive a recommendation, including when a simpler route is the better fit.
€0
30 minutes · written takeaway within 2 business days
Times are shown in your own time zone. We work with clients across time zones.
Describe the process in the form