Why the label "AI company" tells you nothing
The phrase covers a design agency that added a chat widget, a consultancy reselling somebody else's platform, a two-person prompt shop, and an engineering team that has put retrieval, evaluation and integrations into production for paying clients. All four use the same vocabulary on the first call. The vocabulary is free. The capability behind it is not.
What you are actually buying is a set of engineering decisions and the people who will make them: how the system knows what it knows, how anyone will tell whether it is working, where your data goes, and who is responsible when a queue backs up at two in the morning. None of that is visible on a website. It surfaces only when you ask specific questions before the proposal lands.
The nine questions to ask
1. Can they show production systems, not demos?
A demo proves a model can produce plausible output on a curated input, and nothing about the vendor. Ask instead for a system that has run for real users on real traffic for months, then ask what broke and what changed. A team that has operated software answers with specifics: a retrieval index that went stale, a cost spike from an unrouted model call, a tool call that needed an approval step. A team that has only demoed talks about features. Under NDA they can still walk you through the architecture and failure modes without naming the client.
2. Who actually does the work?
Find out who will write the code. Not who leads the account or presents the architecture, but who will be in the repository on a Wednesday afternoon. Ask for names and backgrounds, whether they stay for the duration or rotate off once the contract is signed, and whether any of the build is subcontracted or offshored. Then ask for those engineers to join the next call; if none can before you sign, do not expect one afterwards.
3. How do they evaluate quality?
This question separates engineering teams from everyone else. A language model's output varies, so the only way to know whether a change helped or hurt is to measure it against a fixed set of real examples. Ask how they build that set, who writes the expected answers, and what happens when a score drops. Ask whether retrieval is scored on its own, since fetching the wrong document and misreading the right one are different faults. If the team reads outputs and adjusts the prompt until it looks right, the process cannot be repeated, handed over or defended to a board.
4. How do they handle your data and security?
Start with where your data is stored, where it is processed and which third parties see it. Then the model provider: under the terms actually signed, are prompts and documents retained or used for training? Next, the harder ones: how is one client's data kept apart from another's, how do permissions on your documents carry through into retrieval, how are secrets managed, and what happens about prompt injection when the model reads content it did not write? Get the answers in writing, with a named owner. A team that runs its own infrastructure describes the controls: access policies, encryption, network boundaries, audit logging. A team that leans on a platform describes the brochure.
5. Which models, why, and can you swap them?
Useful opinions on model selection are conditional: this model for that task, a smaller one where it passes the evaluation, an open-weight model where residency or cost demands it. Ask why they chose it and what evidence supports the choice. Then ask what it would take to switch providers. If the answer involves a rewrite, the model is not a component in their architecture. It is the architecture. Providers change prices, retire models and have outages, and your system should survive all three.
6. What does the architecture look like?
Ask them to draw it. Not a slide, but the system they would build for you: where requests enter, what orchestrates the steps, where retrieval sits, where the model is called, how outputs reach your systems of record, and where a human approves an action before it happens. Ask which parts run in your cloud account and which in theirs. A team that has built these systems draws it in the meeting and tells you which boxes you do not need yet. A team that has not promises it in the proposal.
7. How do they price, and what can you exit?
Ask for the pricing model, the stages and what you hold at the end of each. The healthy shape is staged: every stage has a defined output you could take elsewhere, a price agreed before it starts and a clean stopping point after it, with working software arriving early rather than at the end. The unhealthy shapes are an open-ended retainer with no defined outputs, or a large upfront fee with the first working software many months away. Ask what happens if you stop after the first stage.
8. Who owns the IP and the code?
Read the contract for three things. Who owns the code, the prompts, the evaluation data and the infrastructure definitions, and from what point. Whether any part of the system depends on a proprietary platform, framework or licence you cannot take with you. Whether the repositories and cloud accounts are yours or theirs. The right answer is that you own what you pay for from the first invoice, it lives in accounts you control, and the vendor's continued involvement is a choice rather than a dependency.
9. What happens after launch?
An AI system is not finished when it goes live. Models get retired, source documents change, usage shifts and the cost curve moves. Ask who monitors the system, what they watch, how you find out when quality drops, and how changes and fixes are handled. Ask whether documentation and handover are included or invoiced separately, and whether your own team could run the system. A vendor with an operating discipline describes tracing, cost telemetry, alarms and a release process. One without it describes a support email address.
Red flags
Some answers should end the conversation.
- Prompt wrappers. The whole system is a prompt, a model call and a screen. No retrieval, no evaluation, no integration with your data.
- Demo-only portfolios. Every example is a proof of concept or an internal tool with no users. Nothing has met production traffic.
- Platform lock-in. The proposal depends on a proprietary orchestration platform, a no-code builder or the vendor's own hosted service. Leaving means starting again.
- Open-ended retainers. A monthly fee with no defined deliverables and no exit. The incentive is to keep the project going, not to finish it.
- No evaluation story. Quality is asserted rather than measured. Nobody can tell you how they would know if a change made the system worse.
- No engineers on the call. Two meetings in, you have met a founder and an account manager. The builders are a slide.
A scorecard to take into the meeting
Score every vendor against the same nine criteria, so you compare evidence rather than presentation.
| Criterion | What good looks like | What to ask for |
|---|---|---|
| Production evidence | Systems running for real users for months; the team can say what broke and what changed. | A walkthrough of one live system and a call with its owner. |
| People | Named senior engineers who stay on the project; any subcontracting disclosed upfront. | Names, backgrounds and an engineer on the next call. |
| Evaluation | A test set built with your domain experts; regression checks before release; retrieval scored separately. | The evaluation pipeline and an anonymised example from a past project. |
| Data and security | Written answers on storage, processing, third-party access, tenant separation and secrets; controls you can inspect. | A written security summary and a named owner. |
| Models | Choices tied to your requirements and backed by evidence; a provider swap without a rewrite. | The reasoning behind the choice and the cost of switching. |
| Architecture | A diagram they can draw and defend in the room; clear about what runs in your account and what runs in theirs. | A whiteboard session before the proposal arrives. |
| Pricing and exit | Scoped stages with defined outputs, priced before they start; you can stop between any two. | The stage list, the prices and what you keep if you stop. |
| IP and code | You own code, prompts, evaluation data and infrastructure from the first invoice; no proprietary dependency. | The IP clause and a list of every licensed component. |
| After launch | Monitoring, cost telemetry, alarms, documented handover; your own team could run it if you chose to. | The operating model and exactly what handover includes. |
How ComplxAI answers these questions
Since we wrote the scorecard, we should fill it in for ourselves.
ComplxAI is an Australian AI engineering and custom software company: a small senior team that does the work itself, so the engineers who scope a project build it. The engagement runs in five stages: two free conversations, then Discovery, Strategy and Implementation in 12-week blocks, each priced in writing before it begins, with an exit at every boundary. You own the IP from day one of any paid stage, there are no mandatory retainers, documentation and handover are included, and you can take the software in-house. The detail is in how we work.
Examples of what we have shipped, NDA engagements included, are on our case studies page. What we build is described under custom AI development and AI engineering: retrieval, model routing, evaluation, tracing and AWS deployment in the region you need, in your own AWS account or one we set up for you. Put every vendor, including us, through the same nine questions.
Where to go next
- Not sure the roadmap needs AI at all? Start with AI consulting.
- Need a budget range before the first vendor meeting? Use the AI project cost calculator or read the Australian AI development cost report.
- Comparing proposals already? Request a proposal from us and score it against the table above.
- More guides are on the insights page.