Private LLMs under your control
Need to work with confidential documents, cut the cost of repetitive tasks or run AI on your own computer? We select the model and hardware, test it on your examples and deploy the system. Where the task calls for it, we fine-tune a model or train a smaller one for the job.
Working with us
- Who we help
- Businesses and individualsPoland and the EU, including first-time local AI users.
- Where it runs
- Your computer, server or cloudThe environment follows your data and workload.
- What you receive
- A working model and test resultsConfiguration, operating instructions and a maintenance plan.
- Cost
- Quoted after scopingSetup, hardware or hosting, and support priced separately.
When it helps
An employee cannot send a contract to a public chatbot. A document queue is driving up API bills. A model keeps confusing internal categories despite a lengthy prompt. Each problem calls for a different solution.
Where we start
We test one task on an agreed set of examples. We compare quality, response time and total cost against an API service. The results help decide whether you need private hosting, RAG, fine-tuning or a small classifier.
From a first installation to a custom model
- Model selection and evaluation
- We test Polish, domain terminology, difficult examples and missing information. You receive a comparison and acceptance criteria. We also check the specific model licence and compatibility with the serving engine.
- Hosting and deployment
- We deploy on a workstation, an on-premise server or rented infrastructure, including EU regions. The scope covers access controls, an application or API, integrations and monitoring. We select quantization and capacity for the actual workload.
- RAG and fine-tuning
- RAG supplies current document passages to a model. Fine-tuning uses examples to teach a task, such as ticket classification or your required output format. The two can work together. A separate test set shows whether training improved the result.
- Small models for a specific job
- For classification, routing or text scoring, we assess smaller models, distillation and training on labelled data. Training from scratch is scoped after reviewing the dataset and compute budget. Choosing one of several categories rarely needs a large general-purpose model.
Qwen, DeepSeek, Kimi, GLM. Start with the task.
Current candidates include Qwen3.8, DeepSeek-V4.1-Flash, Kimi K3 and GLM-5.3, alongside Mistral Small 4, Gemma 4 and gpt-oss. For Polish, we also assess Bielik and PLLuM. We verify available weights, hardware needs and the exact licence. Open weights gives access to the model; its terms determine permitted use and modification.
An MoE model with relatively few active parameters can still need several GPUs to store all its weights. A family name or leaderboard result does not tell you whether a model fits your hardware.
Compare current models and deployment optionsWhere a private model earns its place
Law firms and legal departments
Search contracts, compare clauses and prepare material for review. Document permissions and source references belong in the application. A lawyer reviews the result.
Manufacturing and engineering
Work with technical documents, service manuals and fault reports. A local installation can run in an isolated network when all its dependencies are available there too.
Accounting, insurance and document operations
Classify correspondence, extract fields and flag missing information. We agree which outputs need approval and measure the cost of errors separately for each document type.
E-commerce, software teams and help desks
Route tickets, search code and draft answers from documentation. For recurring workloads, we compare private inference with API costs, including the cost of corrections.
Want local AI for yourself?
We help individuals run models on computers they already own. We check memory, the GPU and operating system, choose a model, install the tools and show you how to use them. If you are planning a purchase, we recommend a configuration for your task and budget. Renting a GPU can also help test the workload before buying hardware.
Start with your use case and computer specifications, if you have them. You do not need to send private files for the first conversation.
Send your use casePrivacy follows the whole data path
Downloadable weights let you choose where inference runs. We also inspect OCR, embeddings, logs, backups, telemetry and connected tools. We define what may leave the network, who has access and how data is deleted. An EU server alone does not establish GDPR compliance.
When should you use OpenAI, Claude or Gemini?
An API can be more economical for low or irregular traffic, demanding tasks and teams with limited operating capacity. Business services have different data terms from consumer chatbots. We compare specific products and contracts. A hybrid setup is also possible: local tasks alongside agreed API calls with limits on the data sent.
How the work proceeds
- 1
Discuss the task
We establish the expected output, workload, data constraints and available hardware. That gives us a basis for a scoped proposal.
- 2
Test representative examples
We compare models on a separate evaluation set, measuring errors and load. You see the result and cost before choosing a deployment.
- 3
Deploy and hand over
We launch the agreed configuration, connect it to your workflow and provide operating instructions. We agree on updates, recovery and ongoing support.
Jev 1.13 and small decision models
TypeSafe AI’s Jev 1.13, available through OpenRouter, returns structured decisions with probabilities, an interesting approach to routing and classification. We distinguish the Jev service from local projects inspired by it, such as Kev and TinyJev. A valid output format still needs a check that the decision is correct.
Read the introduction to JevBefore you deploy
Do I need my own GPU?
No. We can assess your computer, specify new hardware or rent a GPU. Requirements depend on the model, document length and concurrent workload. Where performance is uncertain, we test before a purchase.Will a local LLM cost less than an API?
That depends on utilization and operating costs. We include hardware or rental, energy, engineering time, recovery capacity and corrections. We compare the cost of a successfully completed task; token charges alone are incomplete.Does a model learn automatically from my documents?
Ordinary inference does not change the model weights. RAG can supply documents at query time. Fine-tuning is a separate training process using prepared examples, access controls and an independent evaluation set.Can you adapt a model to Polish and our industry?
Yes. We first evaluate an existing model on Polish examples from your work. If the errors justify fine-tuning, we prepare training data and compare the model before and after training. We also assess abstention and unfamiliar inputs.Does self-hosting ensure GDPR compliance?
Self-hosting gives you control of the environment. Compliance also depends on the processing basis, permissions, retention and contracts. We provide technical documentation for your data protection reviewers.Do you help individuals without technical experience?
Yes. We can help from computer selection through first use. The scope can include installation, basic privacy settings, instructions and a practical introduction for your intended task.
Private model guides
- Qwen, DeepSeek, Kimi and GLM: what to self-host?
- Local LLM or OpenAI, Claude and Gemini?
- What does a private LLM really cost?
- Fine-tuning, RAG or training your own model?
- How to evaluate an LLM before deployment
- Local LLM on your computer: where to start
- Private LLMs: data, EU hosting and on-premise
- TypeSafe Jev 1.13: a model for application decisions
- Jev 1.13 for business classification and routing
- Local Jev 1.13 alternatives: Kev and TinyJev
Tell us what your model needs to do.
Describe your data and the work that takes the most time today. Add hardware specifications if you have them. We will suggest a practical next step.
- A 30-minute call with the engineer who would lead the work.
- A review of the processes that cost you the most time and money.
- A written summary: a possible direction, missing information and the next step.
The scan chooses one process to assess, and within 2 business days you receive a recommendation, including when a simpler route is the better fit.