Every AI pilot has the same shape. Someone builds a prototype in a week, it answers five questions impressively, and the room agrees this changes everything. Six weeks later the project is quietly paused because the assistant told a customer a delivery date that did not exist.
The failure is architectural, not model choice
Swapping models rarely fixes this. A more capable model is still a model: given a question it cannot answer from context, it will produce the most plausible-sounding continuation. Plausible is not the same as true, and in a customer conversation the difference is a refund.
The fix is structural. Retrieval decides what the model is allowed to see. Guardrails decide what it is allowed to say. Escalation decides what happens when neither is sufficient. Only after those three exist does model selection become the interesting question.
What to specify before writing a prompt
Name the sources of truth and where they live. Decide what the assistant may never state without a retrieved record — price, availability, dates, policy. Write the escalation contract: what the human receives and how fast. Build an evaluation set of real historical questions, including the awkward ones, and gate launch on performance against it rather than on a demo.
None of that is glamorous, and none of it appears in a pitch deck. It is the difference between an assistant you can put in front of customers and one that stays in a sandbox for a year.