xcentity
applied ai lab / agents in production

Agents that show their work.

xcentity runs AI agents inside the tools your team already uses — reading the ticket, moving the data, drafting the next step. Every action is logged, scoped and reversible, so you can hand over real work instead of watching a demo.

run 4812 — intake

Reconcile 1,204 invoices

Triggered by a message in #finance-ops at 08:14.

run 4812 — tool calls
read ledger · 1,204 rows
matched 1,186 against bank feed
18 exceptions queued for review
run 4812 — permissions
one run, three frames

An agent is accountable, or it's a demo.

what we build / 02

Most pilots die in the same place. The model is impressive on a clean example, then nobody can say what it touched, who approved it, or what happens when it is wrong at 2am. Trust never arrives, and the workflow quietly goes back to a person and a spreadsheet.

We build the other half. Each agent gets a narrow scope, credentials your admin controls, a written trace of every call it makes, and a stop button that actually stops it. When it is unsure, it escalates with the evidence attached rather than guessing confidently.

what it touches
Your ticketing, CRM, warehouse, inbox and internal APIs — read-only until you widen it.
what it does
Multi-step work with checkpoints: gather, reconcile, draft, file, and flag what it could not resolve.
what you get
A run log you can audit, an approval queue, cost per run, and the ability to roll any action back.

Anatomy of a single run

pick a step / 03
A task arrives with its context

A message, a webhook or a schedule starts the run. The agent restates the task in one line and names the systems it expects to touch, before it touches anything.

The plan is written down first

Steps, expected cost and the permissions each one needs. Anything outside the agreed scope stops here and asks, rather than improvising a workaround.

Every call is recorded as it happens

Inputs, outputs, latency and cost per call. Results are checked against the source system, so a confident wrong answer shows up as a mismatch instead of a finished task.

A human sees anything irreversible

Payments, deletions, external messages and low-confidence calls queue for approval with the evidence attached. You approve in one click, or send it back with a note.

It ends where your team already works

Results land in the ticket, the sheet or the thread that started it, with a link to the full trace. Exceptions are listed in plain language, not buried in a log file.

Two years in production, measured where it counts.

From customer workspaces, twelve months to September 2026.

0runs completed across 60+ teams
0of steps verified against the source system
0returned per person, per month
0pilots reach production within a quarter

Three days to your first agent

how a pilot runs / 05
day one — map

We watch the workflow as it really is

Ninety minutes over your screen, not a requirements doc. By evening you get the steps written out, with the two we think an agent should own and the ones it should never touch.

day two — shadow

The agent runs beside your team

Read-only, on live work, producing the output it would have filed. You compare its run log against what your team actually did before anything is allowed to write.

day three — ship

Permissions, limits, go live

Scopes set by your admin, spend caps per run, an approval queue for anything irreversible, and a stop button in the dashboard. You keep the keys.

after

Two weeks of tuning, included

We watch the first few hundred runs with you and fix what the real data exposes. Most exceptions turn out to be one missing rule, not a model problem.

Send us the workflow nobody wants.

The repetitive one, held together by a spreadsheet and one person's memory. Tell us what it touches and how often it runs, and we will tell you honestly whether an agent should.

hello@xcentity.ai
two pilot slots open in October