Blog
Your company's own AI agent — where to start
An AI agent is a program with access to your data and permission to perform a specific action. That is all it is. Everything else is engineering: choosing the process, preparing the data, controlling permissions, and measuring results. The most common failure in a first deployment has nothing to do with the model. It happens because nobody in the company can describe the process the agent is meant to take over.
What is the difference between a chatbot, RAG and an AI agent?
A chatbot answers from knowledge baked into the model. RAG adds retrieval over your own documents, so the answer is grounded in specific passages and cites a source. An agent goes further: it holds tools and permissions, so it actually does something. The practical difference is what breaks when the model is wrong.
| Type | Source of knowledge | What it does | Failure mode | Good enough for |
|---|---|---|---|---|
| Chatbot | the model's general training | generates text | invented answer, no audit trail | drafting, translation, brainstorming |
| RAG assistant | your documents, vector index | answers and cites the source | wrong passage, but a human sees the link and can verify | procedures, technical docs, support, internal search |
| Agent with tools | documents plus your system APIs | opens tickets, updates records, drafts documents, sends messages | it changes system state — needs scoped permissions, logs and an undo path | repetitive operations with a clear input and output |
Order matters. Most companies need a RAG assistant first, and an agent only once that assistant has proven, on numbers, that its answers hold up. Reversing the order gives you a system that confidently performs the wrong action.
Which process should the first AI agent handle?
One that meets four conditions at once: it repeats often, it consumes real staff time, it has a clear input and a clear output, and a mistake can be undone. If any of the four is missing, the project usually ends as a demo nobody opens after three weeks.
- Frequency. At least a few dozen occurrences a month. Below that there is nothing to measure and nothing to build a test set from.
- Time cost. A single run takes a person fifteen minutes or more, and that person's time is expensive.
- Clear boundaries. You can state in one sentence what goes in (an e-mail, a PDF, a ticket) and what must come out (a classification, a filled form, an answer with a source link).
- Reversible errors. A misrouted ticket costs a minute. A wrong payment or a binding quote sent to a customer is a different risk class entirely.
- A named owner. Someone who knows the process, has time for reviews, and is allowed to say "this answer is wrong". Without that person there is nothing to improve.
Why "it is all in SharePoint" is not a knowledge base
Because a file share is not a knowledge base. A knowledge base is a set of documents where you know which version is current, who owns it, and how long it stays valid. A typical company drive holds three versions of the same procedure, scans with no text layer, and decisions that live only in chat threads.
Cleaning that up is usually a bigger share of the project than connecting the model. The sequence we follow:
- Inventory the sources. File storage, ticketing system, CRM, mailboxes, wiki, spreadsheets.
- Pick one version of the truth. For each topic, one binding document; the rest is archived. That is a business decision, not a technical one.
- Clean the material. OCR for scans, deduplication, splitting 200-page files into meaningful chunks.
- Add metadata. Date, owner, department, validity. This lets the agent answer "per the March procedure" instead of quoting something from two years ago.
- Mirror permissions in the index. The agent must see exactly what the person asking is allowed to see. Otherwise your assistant becomes the fastest data leak in the building.
One rule we never drop: the agent always cites its source. An answer without a link to the document is worthless, because nobody can check it.
What does "our data never leaves the company" actually mean?
Three different things are sold under that phrase: a vendor's cloud model covered by a data processing agreement, a model running inside your own cloud tenancy, and an open-weights model on a server in your rack. The GDPR does not ban the cloud. It requires a legal basis, a processing agreement, and control over what happens to the data.
A cloud model is usually enough when
- you process ordinary business and B2B contact data, with no special categories involved,
- the vendor signs a DPA, processes within the EEA, and does not train on your data,
- prompt retention is set to zero or to a short, documented window,
- your own customer contracts do not forbid passing their data to third parties.
A local model earns its cost when
- special categories of data are involved, or trade secrets of high value: engineering drawings, formulations, a client's source code,
- a customer or a regulator explicitly requires processing on your own infrastructure,
- the system must work on an air-gapped network — a factory floor, a construction site, a vessel at sea.
A local model is not free just because there is no vendor invoice. You buy and maintain GPU hardware, update the model, watch throughput, and staff it. For classification, data extraction and document question answering, open models are good enough in practice. For long-context reasoning the commercial models still lead by a visible margin.
Either way, prompt logs are personal data and follow the same rules as the rest of the system. Set retention from day one, pseudonymise what you can, and include the agent in your record of processing activities. Obligations under the EU AI Act depend on the use case — an internal knowledge assistant sits in a different category than a system that screens job candidates.
How do you measure whether the agent works?
Instead of "it feels helpful", measure four things: coverage, quality, time and unit cost. There is one precondition — collect the baseline before deployment. Once the agent is live, nobody remembers how long a single case really took.
- Coverage. The share of cases handled without escalation to a human. Rising coverage at constant quality is the only real proof of progress.
- Quality. The share of answers accepted without edits, and the share of answers not backed by a source, scored manually on a random sample.
- Time. Median handling time before and after. Median, not average — outliers distort the picture.
- Cost per case. Model calls plus maintenance, divided by cases handled. This number decides whether scaling makes sense.
Add a regression set: 50 to 100 real questions with approved answers, run after every change to the prompt, the model or the knowledge base. Without it you cannot tell whether a change improved things or quietly broke an area nobody checked.
When is an AI agent the wrong answer?
When the same result is cheaper and more reliable without a model. We say so even when the client has already decided.
- The process is not written down. If five people do it five ways, the agent will freeze one of them at random. Describe first, automate second.
- The task is deterministic. Format validation, copying a field between systems, looking up a rate in a table — an API integration or a rule handles it. A model only adds uncertainty and cost.
- You need exact repeatability. Settlements, legal documents, machine control. Language models are probabilistic by design.
- The decision cannot be undone. Payments, terminations, binding quotes. Here the agent may only prepare a draft for a human to approve.
- Nobody wants to clean the data. Building on a mess produces a fast demo and a lasting disappointment.
Sometimes the best "AI project" is an integration between two systems that removes the question entirely.
Where to start
With one process, one dataset and one metric. A week to pick the process and record the baseline, then a knowledge assistant on real documents, and only then permissions to act — granted one at a time, logged, and reversible. Keep a human approving until the numbers say otherwise.
We build these systems the same way we build our own products — the BARVEA BIM/CDE platform, the ClimaBox IoT device, the autoPolar sailing trim assistant — with the person who maintains them a year from now in mind. The scope of work is in our services, and the systems themselves in projects.
We do not build template websites, sell hosting without a project, or rent out developers to other people's teams. If you have a process that looks like a candidate for automation, get in touch — after the call we come back with a quote within one working day.