These 50 questions are grouped into five architecture domains. None of them are meant to have an easy yes — a lot of "no" or "not yet" answers is normal for a pilot. The point is knowing which ones, not discovering them in production.

01–10 · Data & Integration Architecture

  1. Where does the system of record live for every data source this AI system touches — and does the AI system ever become a second source of truth by accident?
  2. What is the data freshness requirement, and does your ingestion cadence actually meet it?
  3. Is there a documented data lineage from source system to the point where a model or agent consumes it?
  4. What happens when a source system is unavailable — does the AI system fail loudly, or silently serve stale data?
  5. Are documents and records classified by sensitivity before they're indexed, or after?
  6. Who owns each data source feeding the system, and do they know it's being used this way?
  7. Is there a process for removing data from the index when it's deleted or revoked at the source?
  8. Does the architecture support incremental updates, or does every change require a full re-index?
  9. What is the maximum acceptable latency between a source-system change and that change reaching AI outputs?
  10. Have you tested what happens when two data sources disagree?

11–20 · AI / Model Architecture

  1. Is the choice of RAG, fine-tuning, or agentic architecture driven by the actual problem — or by which one is currently popular?
  2. If using RAG, is retrieval hybrid (lexical + vector), or vector-only?
  3. Are generated answers grounded with citations back to source content, or asserted from model memory?
  4. What is the fallback behavior when retrieval returns nothing relevant?
  5. Is there a defined process for evaluating a model or prompt change before it reaches production?
  6. Is the system built around a single model, or does it route between models by task difficulty?
  7. What happens when the model provider changes pricing, deprecates a model, or has an outage?
  8. Is context-window usage bounded, or can a single request silently consume unbounded tokens?
  9. Is there a documented reason agentic autonomy was chosen over a simpler fixed workflow?
  10. Can the system's reasoning and output be audited after the fact, not just observed live?

21–30 · Security

  1. Is authorization enforced at the point of retrieval, or only at the application UI layer?
  2. Can a user, through the AI system, ever retrieve content they wouldn't be authorized to see directly?
  3. Are tool and agent permissions scoped per action, or does the agent hold standing broad credentials?
  4. Is there a defined threat model for this specific AI system, not just a generic security review?
  5. How is prompt-injection risk handled for any input the system treats as instructions?
  6. Is there logging sufficient to reconstruct what data was retrieved, and why, after an incident?
  7. Are secrets and credentials used by agents or tools rotated and scoped the same as any other system credential?
  8. What is the blast radius if this AI system is compromised — what else could it reach?
  9. Has the system been tested against malformed or adversarial inputs, not just happy-path inputs?
  10. Is there a kill switch — a fast, tested way to disable the system if it misbehaves in production?

31–40 · Governance & Responsible AI

  1. Is this AI system in a central inventory, or does its existence depend on someone remembering to mention it?
  2. Is there a named owner accountable for this system's behavior in production?
  3. Is the level of governance applied proportional to the system's actual risk tier — or applied uniformly regardless of risk?
  4. Is there a human-approval checkpoint for any action with real-world consequence — financial, legal, or irreversible?
  5. Is there a documented policy for what the system is, and isn't, allowed to do?
  6. Who is notified when the system behaves unexpectedly, and what is the expected response time?
  7. Is there a process for periodically re-reviewing the system as usage or data changes?
  8. If regulation changes, is there a way to know which systems are affected?
  9. Is bias or fairness evaluation relevant to this use case — and if so, has it actually been done?
  10. Can you explain, in plain language, why the system produced a specific output, if asked?

41–50 · Production & Operations

  1. Is there a defined SLA or latency budget, and does the architecture meet it under real load, not demo load?
  2. Is cost per request and cost per user actually measured, or only estimated once at design time?
  3. Is there a caching strategy, and has it been tested for staleness risk?
  4. What is the rollback plan if a new model version or prompt change degrades quality?
  5. Are query, retrieval, and authorization events logged as separate, distinguishable signals?
  6. Is there a standing evaluation set re-run on every material change, not just at launch?
  7. Does the system have defined rate limits, and what happens to a user who exceeds them?
  8. Who is on call when this system fails, and do they have the context to debug it?
  9. Has the system been load-tested at realistic peak volume, not average volume?
  10. If this system succeeded far beyond expectations tomorrow, what would break first?

AQEVON's point of view

None of these questions require exotic technology to answer well. Almost every "no" on this list traces back to the same root cause: a pilot's architecture was never revisited once it started working. Production-readiness is a day-one architecture decision, not a pre-launch checklist bolted on afterward — this list is most useful used early, not as a gate right before launch.