SERVICES / LLM AND AI AGENTS
Enterprise LLM and AI Agent Development
Retrieval systems, agentic workflows, and model fine-tuning built to run inside your infrastructure. On your hardware or ours, never on a third-party endpoint.
WHAT WE BUILD
Core LLM & AI Agentic Capabilities
Private RAG systems
Document ingestion, chunking, embedding, and retrieval over your own corpus. The vector store sits on your infrastructure. No document reaches an external API at any point in the pipeline.
Agentic workflows
Multi-step agents that read, decide, and act against your systems, with a human approval gate on every write operation. Built for the processes that currently consume a person’s afternoon.
Fine-tuning on customer infrastructure
Domain adaptation on your data, on your hardware. Your acronyms, your document taxonomy, your schema formats. The training data and the resulting model stay with you.
MCP integration
Model Context Protocol connectors between language models and the systems you already run, ERP, CRM, ticketing, document management. Including for models and hardware we did not supply.
How We Build It
Open weights, on your terms
We deploy open-weight models under permissive licences, selected per workload rather than by default. No dependency on a single vendor’s API, no repricing risk, no model deprecated out from under a production system.
Approval gates on every write
Agents prepare transactions; a person commits them. This is architectural in everything we build, not a configuration option that can be switched off under delivery pressure.
Deployed where your data is allowed to be
On-premise, air-gapped, or single-tenant hosted in a region you choose. The deployment target is a design input from the first week, not a port at the end.
Built by a team that ships AI into regulated environments
ARSA has been building and operating AI systems in production since 2018, across defence, law enforcement, and industrial deployments where the system has to keep working and the data cannot leave the building. The same engineers build the language-model systems described on this page.
Every engagement starts with a paid feasibility assessment: an operational diagnosis, a technical feasibility judgement against your actual data and infrastructure, and a cost model. Two weeks. We will tell you if it should not be built.
TIERS
Engagement
| Stage | From | Duration |
|---|---|---|
| Feasibility assessment | $4,500 | 2 weeks |
| Pilot deployment | $20,000 | 8 weeks |
| Production programme | $60,000-$300,000+ | 12 weeks+ |
The feasibility assessment fee is deducted from the project fee if you contract within 90 days.
FAQ
AI LLM & RAG Service Questions
Which models do you deploy?
Open-weight models under permissive licences, selected per workload, currently the Qwen and Llama families for general reasoning, plus specialist models for OCR and embedding. Model choice is a design decision made in the assessment, not a default.
Does our data leave our infrastructure?
No. Retrieval, inference, and fine-tuning all run on infrastructure you control. Where we host on your behalf, it is a single-tenant instance in a region you choose.
Who owns a fine-tuned model?
Your data remains yours in every case. A model adapted exclusively on your data for your use case is typically yours to use; the base weights remain under their original open-weight licence.
Can you integrate with systems we already run?
Yes. MCP connectors for ERP, CRM, ticketing, and document systems, including where the model and the hardware were not supplied by ARSA.
How do you stop an agent doing something irreversible?
Hard-coded human-in-the-loop approval on every write operation. Agents prepare transactions; a person commits them. This is architectural, not configurable.
What if we already run models on cloud APIs?
Then the question is whether you need to stop, and why. Data residency, cost at scale, and vendor deprecation risk are the three reasons clients move. If none of them apply to you, cloud is a reasonable place to stay and we will say so.
