Quick answer: Evaluate an AI creative or automation agency on five dimensions — capability proof, governance maturity, security posture, workflow transparency, and commercial structure — using written answers, live work and a bounded pilot. The twelve RFP questions below cover all five; an agency ready for enterprise work answers them in writing within days, not weeks.
Enterprise procurement teams know how to evaluate software vendors and how to evaluate creative agencies. AI creative and automation agencies are both at once, which is why standard templates miss the risks that matter. This guide gives procurement and marketing leaders a complete evaluation structure: the dimensions, the questions, the red flags, and the pilot design that converts a promising pitch into verified capability.
The five evaluation dimensions
- Capability proof. Live, recent, attributable work — not a showreel of cherry-picked frames. Ask to see work in your category, and ask what part of it was AI-produced versus conventionally produced.
- Governance maturity. A written responsible-AI practice: disclosure policy, human accountability gates, provenance capability, talent-consent handling. Singapore's IMDA frameworks give you the vocabulary to test against — see our responsible AI explainer.
- Security posture. Which tools at which tiers, no-training terms, access control, retention. The ten-question checklist in our AI security guide is the instrument.
- Workflow transparency. The agency can walk you through its production system stage by stage, including where humans sign. A capable agency has a workflow it can draw on a whiteboard; a tool-user has a subscription list.
- Commercial structure. Bounded pilot, clear per-band pricing, and terms that keep your assets, references and any fine-tuned artefacts yours.
The twelve RFP questions
Grouped by dimension, with what a strong answer looks like:
| Question | Strong answer looks like |
|---|---|
| 1. Show three recent projects comparable to our scope, with what was AI-produced. | Live links or full assets, honest AI/conventional breakdown, results where measurable |
| 2. How do you keep hundreds of assets consistent with our brand system? | Reference-locking method described concretely; a no-regeneration rule for approved assets |
| 3. What can't AI production do well for our category today? | A specific, unhesitating answer — honesty here predicts honesty everywhere |
| 4. Who is the named human accountable for each published asset? | Named gates: creative QC owner, compliance owner, final sign-off |
| 5. Show us your disclosure policy for AI-generated content. | A written, tiered policy — not "we can discuss this" |
| 6. Can your outputs carry C2PA Content Credentials, and how do you handle talent likeness consent? | Yes on provenance capability; consent process described with AI-use coverage |
| 7. Which AI tools and models will touch our material, at which tiers? | A named list with enterprise/API tiers and no-training terms |
| 8. What are your data retention and deletion terms, including prompts and rejected generations? | A stated schedule covering production residue, with deletion confirmation |
| 9. How do you prevent unapproved AI tool use on our account? | An approved-tool registry and staff policy, reviewed on a schedule |
| 10. Walk us through your workflow from brief to delivery, with review structure. | Staged workflow with consolidated annotation rounds and an agreed fix-round structure |
| 11. What does a pilot look like, and what will it cost? | Bounded scope, agreed success metrics, fixed price, 4–6 weeks |
| 12. Who owns the assets, references and any custom artefacts created for us? | You do — stated plainly in the terms |
Red flags that end evaluations
- "Our tools are proprietary / confidential." Model and tier transparency is a security requirement, not a trade secret. Pipelines can be proprietary; data handling cannot be opaque.
- No named humans in the workflow. If nobody signs the gates, your brand is being reviewed by a queue.
- Guarantees of specific creative or ranking outcomes on a date. Honest vendors commit to process, capacity and measured improvement — not clairvoyance.
- Demo-only proof. Everything looks perfect and nothing is live, recent or attributable.
- No answer to question 3. An agency that claims AI does everything well has not shipped enough AI work.
- Pressure to skip the pilot. The pilot protects both sides; resisting it says the capability may not survive one.
Designing the pilot
A good pilot is small enough to be safe and real enough to be predictive. Scope one brand and one defined asset set; agree success metrics before generation starts — QC pass rate, approval-round count, cost per approved asset, cycle time; and run 4–6 weeks. In Singapore, defined-scope enterprise pilots typically start from around S$5,000, with ongoing engines running from about S$5,000 to S$20,000+ a month depending on volume and markets. Evaluate the pilot on the system as much as the assets: did the workflow hold, did the gates function, did review rounds stay consolidated? Assets tell you what the agency can make; the system tells you what year two looks like.
A simple scoring matrix. Weight the five dimensions for your context — a regulated enterprise might run capability 30%, governance 25%, security 25%, workflow 10%, commercial 10% — score each vendor's written answers 1–5 per dimension, and let the pilot decide between the top two. The discipline matters more than the weights: written answers, same questions to every vendor, scores before demos. It keeps the evaluation about the system rather than the showreel.
Frequently Asked Questions
Should we run this as a formal RFP or an informal evaluation?
Match the stakes. A multi-market production engine justifies a formal RFP with the twelve questions as the written round. A single-campaign engagement can run the same questions informally — the point is written answers and a bounded pilot, not the ceremony around them.
How is evaluating an AI agency different from evaluating a traditional agency?
Two dimensions are new: security posture (which AI systems see your data and under what terms) and governance maturity (accountability, disclosure, provenance). Capability proof also changes shape — you are evaluating a production system's consistency at volume, not a team's taste on a dozen hero assets. Workflow and commercial evaluation stay familiar.
What if a vendor scores well on capability but poorly on governance?
Decide whether the gap is maturity or indifference. A capable vendor without written policies can often produce them during the evaluation — that is a good sign, and your questions did them a favour. A vendor who argues the policies are unnecessary is telling you how they will respond to every future compliance request.
Does this framework apply to AI automation agencies too?
Yes, with one addition: for automation and agent work, governance weight increases, because agents act autonomously. IMDA's Model AI Governance Framework for Agentic AI (January 2026) is the reference — ask how the vendor keeps humans accountable for agent outcomes, and how agent actions are logged and reviewable.
How does AI Studio perform against these twelve questions?
They are drawn from what enterprise clients ask us, so we answer all twelve in writing during any evaluation — tool and tier list, disclosure policy, named gates, retention terms, pilot structure and asset ownership included. We would rather be scored on this framework than pitched around it.