Skip to content

Private LLMs under your control

Need to work with confidential documents, cut the cost of repetitive tasks or run AI on your own computer? We select the model and hardware, test it on your examples and deploy the system. Where the task calls for it, we fine-tune a model or train a smaller one for the job.

Working with us

Who we help
Businesses and individualsPoland and the EU, including first-time local AI users.
Where it runs
Your computer, server or cloudThe environment follows your data and workload.
What you receive
A working model and test resultsConfiguration, operating instructions and a maintenance plan.
Cost
Quoted after scopingSetup, hardware or hosting, and support priced separately.

When it helps

An employee cannot send a contract to a public chatbot. A document queue is driving up API bills. A model keeps confusing internal categories despite a lengthy prompt. Each problem calls for a different solution.

Where we start

We test one task on an agreed set of examples. We compare quality, response time and total cost against an API service. The results help decide whether you need private hosting, RAG, fine-tuning or a small classifier.

From a first installation to a custom model

Model selection and evaluation
We test Polish, domain terminology, difficult examples and missing information. You receive a comparison and acceptance criteria. We also check the specific model licence and compatibility with the serving engine.
Hosting and deployment
We deploy on a workstation, an on-premise server or rented infrastructure, including EU regions. The scope covers access controls, an application or API, integrations and monitoring. We select quantization and capacity for the actual workload.
RAG and fine-tuning
RAG supplies current document passages to a model. Fine-tuning uses examples to teach a task, such as ticket classification or your required output format. The two can work together. A separate test set shows whether training improved the result.
Small models for a specific job
For classification, routing or text scoring, we assess smaller models, distillation and training on labelled data. Training from scratch is scoped after reviewing the dataset and compute budget. Choosing one of several categories rarely needs a large general-purpose model.

Qwen, DeepSeek, Kimi, GLM. Start with the task.

Current candidates include Qwen3.8, DeepSeek-V4.1-Flash, Kimi K3 and GLM-5.3, alongside Mistral Small 4, Gemma 4 and gpt-oss. For Polish, we also assess Bielik and PLLuM. We verify available weights, hardware needs and the exact licence. Open weights gives access to the model; its terms determine permitted use and modification.

An MoE model with relatively few active parameters can still need several GPUs to store all its weights. A family name or leaderboard result does not tell you whether a model fits your hardware.

Compare current models and deployment options

Where a private model earns its place

Law firms and legal departments

Search contracts, compare clauses and prepare material for review. Document permissions and source references belong in the application. A lawyer reviews the result.

Manufacturing and engineering

Work with technical documents, service manuals and fault reports. A local installation can run in an isolated network when all its dependencies are available there too.

Accounting, insurance and document operations

Classify correspondence, extract fields and flag missing information. We agree which outputs need approval and measure the cost of errors separately for each document type.

E-commerce, software teams and help desks

Route tickets, search code and draft answers from documentation. For recurring workloads, we compare private inference with API costs, including the cost of corrections.

Want local AI for yourself?

We help individuals run models on computers they already own. We check memory, the GPU and operating system, choose a model, install the tools and show you how to use them. If you are planning a purchase, we recommend a configuration for your task and budget. Renting a GPU can also help test the workload before buying hardware.

Start with your use case and computer specifications, if you have them. You do not need to send private files for the first conversation.

Send your use case

Privacy follows the whole data path

Downloadable weights let you choose where inference runs. We also inspect OCR, embeddings, logs, backups, telemetry and connected tools. We define what may leave the network, who has access and how data is deleted. An EU server alone does not establish GDPR compliance.

When should you use OpenAI, Claude or Gemini?

An API can be more economical for low or irregular traffic, demanding tasks and teams with limited operating capacity. Business services have different data terms from consumer chatbots. We compare specific products and contracts. A hybrid setup is also possible: local tasks alongside agreed API calls with limits on the data sent.

How the work proceeds

  1. 1

    Discuss the task

    We establish the expected output, workload, data constraints and available hardware. That gives us a basis for a scoped proposal.

  2. 2

    Test representative examples

    We compare models on a separate evaluation set, measuring errors and load. You see the result and cost before choosing a deployment.

  3. 3

    Deploy and hand over

    We launch the agreed configuration, connect it to your workflow and provide operating instructions. We agree on updates, recovery and ongoing support.

Jev 1.13 and small decision models

TypeSafe AI’s Jev 1.13, available through OpenRouter, returns structured decisions with probabilities, an interesting approach to routing and classification. We distinguish the Jev service from local projects inspired by it, such as Kev and TinyJev. A valid output format still needs a check that the decision is correct.

Read the introduction to Jev

Before you deploy

  • Do I need my own GPU?

  • Will a local LLM cost less than an API?

  • Does a model learn automatically from my documents?

  • Can you adapt a model to Polish and our industry?

  • Does self-hosting ensure GDPR compliance?

  • Do you help individuals without technical experience?

Tell us what your model needs to do.

Describe your data and the work that takes the most time today. Add hardware specifications if you have them. We will suggest a practical next step.

  • A 30-minute call with the engineer who would lead the work.
  • A review of the processes that cost you the most time and money.
  • A written summary: a possible direction, missing information and the next step.
€030 minutes · written takeaway within 2 business days
Send your use case

The scan chooses one process to assess, and within 2 business days you receive a recommendation, including when a simpler route is the better fit.