A demo is a performance. It is rehearsed, it runs on chosen data, and it exists to compress your scepticism into a signature. None of that makes vendors dishonest — a good demo of a good system is a legitimate thing — but it does mean the demo cannot be the evidence. MIT's 2025 research on enterprise AI found the failed 95% of pilots clustering around exactly this gap: tools that performed in controlled demonstrations and then could not hold context, adapt, or integrate in production.
The correction is not cynicism. It is a better set of questions. The ten below are ordered roughly as a sales conversation unfolds, and for each we describe the answer a serious implementation partner gives — and the answer that should end the meeting. Bring the list. A vendor who resents it has told you something more useful than the demo did.
"What is the written success criterion, and what happens if we miss it?"
The single most diagnostic question in the field. A serious partner answers with a number, a measurement method, a time window, and a consequence — reduced fees, remediation, an exit clause. A demo agency answers with capabilities. If the outcome is not written down, you are buying effort, not a result.
"What does this process cost us today, and how will you measure that before building?"
The right answer describes an audit: process mapping, baseline measurement, a document you keep whether or not you proceed. The wrong answer skips straight to the build — which means the "improvement" will be measured against nothing, and therefore unfalsifiable.
"Show me this working inside a system like ours — not a demo environment."
The demo runs on curated data in a clean sandbox. Your business runs on a specific ERP, a specific accounting stack, and fifteen inconsistent supplier formats. Ask what the integration path into your systems of record looks like, named. A tool that lives in a separate tab is a tool your team will abandon by the first busy week.
"What happens when the system is unsure or wrong?"
Every automated system is sometimes wrong. A serious answer describes confidence thresholds, human review queues, fallback paths, and how errors are detected and corrected. An answer that amounts to "it's very accurate" means error handling has not been designed — and it will be designed later, by you, in production.
"Who maintains this after go-live, at what cost, and what breaks it?"
APIs change, formats drift, models get updated. The honest answer includes a recurring cost and names the fragile points. A proposal with no maintenance line has not eliminated maintenance; it has scheduled it as an emergency.
"Which of your clients can I call — including one where things went sideways?"
References chosen by the vendor prove little; every vendor has two happy clients. The revealing request is the second half. A partner confident in how they handle problems will connect you to a client whose project hit turbulence. A vendor for whom nothing has ever gone wrong is either brand new or editing.
"Who exactly will work on our project, and how many implementations like ours have they personally done?"
The people in the sales meeting are often not the people who deliver. Ask for names. In a field this young, "implementations personally completed" is a fairer credential than any logo wall — and note that in a market growing this fast, most agencies are younger than the invoices they process.
"What do you need from our team, in hours, and when?"
The seductive answer — "nothing, we handle everything" — is the wrong one. Real integration requires your people: documenting the process, testing edge cases, adjudicating exceptions. A partner who quantifies your internal effort is planning a real project. A vendor who waives it is planning a handover of blame.
"What happens to our data — where does it go, what is it used for, and what does the AI Act require of us as deployers?"
The answer should name where data is processed, whether it trains anyone's models, and what your obligations as a deployer look like — without either dismissing the regulation or weaponising it into fear. A vendor who cannot discuss the EU AI Act's actual requirements in 2026 is outsourcing your compliance exposure back to you, silently.
"What should we not automate?"
The final filter, and the most human one. A partner with judgement will name processes in your business where automation is premature — undocumented workflows, judgement-heavy calls, places where the baseline is unmeasured. A vendor for whom everything is automatable is not selling a service; they are selling everything, which is the same as understanding nothing.
What the ten questions have in common
Read back, the list is one question wearing ten costumes: does this vendor make claims that can be checked? Written criteria, baselines, named integrations, error paths, maintenance costs, warts-and-all references, named people, quantified effort, compliance specifics, and declared limits — every item converts a promise into something falsifiable. The MIT research found user trust collapsing fastest around tools whose promised results were never verified.
One closing note of fairness: a vendor who answers all ten well has earned something from you too — a mapped process, honest data access, and internal time actually committed. Verification is a two-way discipline. That is what separates a purchase from a partnership.
Frequently asked questions
What are the biggest red flags when choosing an AI automation vendor? No written success criterion, no baseline measurement before building, no maintenance cost in the proposal, references only from hand-picked happy clients, and the claim that your team will need to invest no time. Each converts later into a project that cannot be evaluated or sustained.
Should an AI vendor guarantee results? A serious partner commits to a written, measurable success criterion with consequences for missing it. Blanket guarantees of specific ROI before your process has been mapped and measured are a sales device, not a commitment.
How do I evaluate an AI vendor demo? Treat the demo as a claim, not evidence. Ask to see the workflow operating on data and systems like yours, ask what happens when the system is unsure, and ask which integration points into your existing software are included in the quoted price.
What questions reveal whether an agency can actually deliver? Ask who personally will do the work and how many similar implementations they have completed, ask to speak to a client whose project had problems, and ask what they would refuse to automate in your business. Specific answers indicate experience; universal enthusiasm indicates marketing.
Sources: MIT NANDA, "The GenAI Divide: State of AI in Business 2025" (2025), on demo-to-production failure patterns and trust erosion around unverified claims. The ten questions are Arkhon's own methodology; we welcome being asked all of them.



