The pilot was fine.The scoping never happened.
A demo that impressed everyone and then quietly stopped being opened is the most common outcome in this category. It is almost never a model problem.
Open source projects
Public on github.com/JIGGAI
AI plugins
Published on npm, free to install
AI plugin installs
66,095 to date, per npm
AI monthly plugin installs
2,811 in August 2026, per npm
What we hear in the first hour.
Not a generic list. These are the ones that come up in businesses shaped like yours.
It worked in the demo. It does not survive a real week.
Demos are built on the happy path. The exceptions are where the work actually is, and nobody mapped them first.
We could not tell whether it was working, so it quietly stopped mattering.
No success measure was defined before the build, so there was nothing to report and nothing to defend at budget time.
The team went back to the spreadsheet within a month.
The tool asked people to change how they work without removing anything from their day. The spreadsheet always wins that trade.
The person who championed it left, and it went with them.
It was never wired into a process — it was a side project with an owner, which is a different and more fragile thing.
Now nobody here wants to hear the word again.
Fair, and useful. A second attempt has to earn trust with something small, measurable and boring.
Pilots rarely fail at the model. They fail at deciding what the model was for.
Which is why the first thing we produce is not a build. It is a ranked list with the reasoning attached — including the things we would tell you not to touch.
The software already in the building.
We do not arrive with a platform to sell you. We arrive expecting these, and the work is usually in the gaps between them.
The pilot itself
Whatever was built. We read it, run it, and work out whether the idea was wrong or only the scoping was.
The spreadsheet it lost to
Usually the most honest documentation of the real process in the building. We start here more often than not.
Your ticketing or CRM
Jira, Linear, HubSpot, Salesforce — wherever work is tracked, and where the gap between the tracked process and the real one shows up.
Whatever the vendor left behind
Prompts, config, a hosted account nobody can log into. We inventory it so you know what is safe to switch off.
The shadow tooling
Group chats, personal scripts, an inbox rule someone wrote in 2019. Undocumented and load-bearing.
Your identity and access setup
Because the honest constraint on any agent is what it is allowed to touch, and that answer lives here.
Restarting on the process that already has a number attached.
Roughly what one entry in your report looks like — a real shape, with the numbers changed.
- What we saw
- A pilot that summarised inbound enquiries but sat beside the workflow rather than inside it. The team still opened every enquiry to check the summary, so it added a step instead of removing one.
- What it costs today
- Time lost across the team re-reading what had already been summarised, plus the ongoing cost of a tool nobody trusts enough to act on.
- What we would build
- The same summarisation, moved into the step where the enquiry is triaged rather than beside it, with the confidence score visible and a one-click path to the original. Nothing is hidden from the person deciding.
- How you would know it worked
- Median time from enquiry arriving to being assigned, and the share of enquiries where someone opened the original anyway. If that share does not fall, the tool is still not trusted and we say so.
The reply itself. Enquiries at this stage are the first impression of your business, and a drafted reply that reads as generated costs more than it saves. The judgement call about what to do stays human — we automate everything up to it.
What coordination work costs you.
Three numbers you already half-know. Move them until they look like your business.
re-keying, chasing status, producing the same document again
on that work specifically, not their whole job
salary, tax, benefits, desk
A 75% capture rate. The other twenty-five per cent is judgement, exceptions, and not wanting to look greedy.
This is arithmetic, not a finding. It rests on three numbers you guessed. The assessment replaces all three with numbers we observed — and tells you which of those hours are actually worth automating.
See what a real finding looks like →What the six days look like.
The most common question we get is not about AI. It is what these people will actually do in my building.
- Day 1Walk the floorWhoever is on shift
We start where the work happens, not in a meeting room. Nobody prepares anything, and the first day is mostly watching.
- Day 2Sit with the people doing itOps, admin, finance
Conversations with the roles that touch the work most, and a real task followed end to end — including the parts that happen in a group chat.
- Day 3Systems, then a read-backWhoever holds the logins
What you run, what talks to what, and where a person is currently the integration. We tell you what we saw before we leave, while it is still cheap to correct.
- Days 4–6Research and discoveryOur desks, not yours
Away from your building. We cost the work we watched, model the alternatives, and test the shortlist against your own numbers rather than a framework.
- +1 weekOne recommendationPresented in person
The single change worth making first, specified precisely enough to build — with the ranked analysis behind it and the list of what we would not automate.
The questions this raises.
We already spent money on this once. Why would we spend again?
Because the first spend bought a build and this one buys a decision. If the assessment concludes that nothing here is worth automating yet, that is a finding you can act on, and it costs a fraction of another project budget.
Will you tell us the last vendor was wrong?
Only if they were, and only with reasoning you can check. More often the build was competent and pointed at the wrong process, which is a scoping failure rather than an engineering one.
Can you fix what we already have instead of starting over?
Often, yes. Re-siting a working component inside the process it should have been in is usually cheaper than rebuilding, and it is the first thing we look for.
Our team is sceptical now. Does that make this harder?
It makes it easier. Sceptical teams tell you where the previous attempt actually broke, which is the information the assessment needs. Enthusiastic teams tend to describe the version they hoped for.
Start with half an hour.
We'll tell you whether an assessment would pay for itself in an operation like yours. Sometimes the answer is no.