September 2026 · 5 min read
The Dogfood Test for Choosing an AI Consultancy

Key Definitions
Dogfood test A vendor due-diligence standard: does the firm run what it sells, at scale, on itself first? For consultancies: does it run agents internally, and can it show internal data to prove it?
Copilot-class agent An agent that assists humans with single-step work such as search, synthesis, and drafting, while humans keep decision and execution. McKinsey’s self-reported 1.5 million hours saved are this class (search and synthesis) — a different capability from end-to-end agents that execute complete workflows.
AI high performer McKinsey’s definition: an organization attributing 5% or more of its EBIT to AI use and describing the impact as significant. Only 6% of respondents qualified in the 2026 global survey — flat year over year.
When choosing an AI consultancy, first ask whether it runs agents on itself — and then what data it can show to prove it. McKinsey runs 25,000 agents beside 40,000 people, EY built 50,000 agents in nine months, and Wipro sells we are our own reference. The consulting industry is now dogfooding at scale — yet the numbers are largely self-reported, and McKinsey’s own survey finds only 6% of firms attribute 5%+ of EBIT to AI. Here is the dogfood test as an executable due-diligence standard.
1. The facts: consultancies moved from selling AI advice to running agent fleets
McKinsey CEO Bob Sternfels (HBR IdeaCast): roughly 25,000 personalized agents running beside 40,000 employees, targeting near 1:1 by year-end; 1.5 million hours saved in search and synthesis work last year; 2.5 million charts generated in six months — a figure not independently verified. Sources: Yahoo Finance and TechTimes.
EY: per its official case study, more than 50,000 agents built in nine months; EY.ai EYQ deployed to 300,000+ professionals; the platform co-built with Microsoft (orchestration) and Nvidia; EY says it drinks its own champagne before sharing what it learns with clients. Source: EY’s official case study.
Wipro: CIO Kenny Kesar, in a Google Cloud Consulting video interview: we are our own reference — an $11 billion services firm incubates and validates solutions internally before selling them; 90 orchestrated agents already run across 15,000 associates on Gemini Enterprise. Source: Wipro/Google video interview summary.
2. The irony: McKinsey’s own survey says only 6% qualify
McKinsey’s 2026 State of AI global survey (1,719 respondents, May 4–June 8, 2026): 88% of organizations regularly use AI in at least one function, but only 39% attribute any level of EBIT impact to AI (most below 5%), and only 6% qualify as AI high performers (5%+ EBIT attribution with significant impact) — flat versus 2025. Source: McKinsey’s official survey page.
In other words: across the industry — consultancies included — most AI spending has not yet become enterprise-level EBIT. And the flagship internal number consultancies sell with, the 1.5 million hours, is copilot-class work (search and synthesis), not end-to-end autonomous agent work.
3. Our judgment: the dogfood test has three layers
Layer one, quantity: how many agents does the firm run internally? A consultancy with no internal agents has no business selling agent programs — it has never run one on itself. Layer two, class: what class of work do its agents do? Copilot-class (search, synthesis, drafting) and end-to-end (executing complete workflows and handling exceptions) are different capabilities; an hours-saved figure without a class breakdown is not a number. Layer three, audit: has anyone independently verified its internal numbers? Self-reported numbers are the wish for evidence; independent audit is evidence.
Our judgment: dogfooding is necessary but not sufficient. It proves the firm is not selling something it would not run itself — it does not prove what it runs will reproduce in your organization. That depends on your processes and data, not on any template it sells.
4. Buyer checklist: five due-diligence questions for an AI consultancy
① How many agents do you run internally, in which functions, for how long?
A consultancy with no internal agents should not be selling agent programs.
② Are your agents copilot-class or end-to-end?
Can it state the share completed autonomously with no human takeover? The two classes are different capabilities; unclassified numbers are not numbers.
③ How many hours did you save — on what baseline, defined by what metric?
Time figures without a baseline cannot be tested; demand the window, sample, and comparison object.
④ What are your error rates and escalation paths?
How many agents have gone out of parameters, and were they found after the fact or in real time? Specific answers are more credible.
⑤ Are the internal numbers third-party audited?
If not, what can it open for your independent verification? Self-reported numbers get the vendor-claim discount.
5. Action steps and the question left for buyers
Next 30 days: put the five questions into your RFP or diligence checklist; require candidate consultancies to show internal agent telemetry (count, class, baseline, error rate, audit status); demand a metric definition for every efficiency X% claim; treat openness to independent verification as a plus and case-studies-only as a risk flag.
Question for buyers: how many of the five questions can your current AI consultancy answer? If it does not run agents on itself, why would you trust it to run them well for you?
OOMeta AI
OOMeta’s position and practice: the dogfood test is how we operate — we only deliver what we have run internally, and we accept client diligence through internal telemetry and independent verification. This article turns that standard into a reusable buyer checklist, to push case-study-only selling out of the market.
Schedule a DiagnosticReferences: McKinsey, The State of AI: Global Survey 2026 https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai ; Yahoo Finance (McKinsey CEO: 25,000 agents / 1.5M hours) https://finance.yahoo.com/news/mckinseys-ceo-breaks-down-ai-100301404.html ; TechTimes (McKinsey 2026 survey: 6% high performers; 2.5M charts unverified) https://www.techtimes.com/articles/325590/20260826/record-ai-spending-cant-move-earnings-needle-94-enterprises-mckinsey-finds.htm ; EY, Building an enterprise-scale agentic AI operating system https://www.ey.com/en_gl/insights/ai/building-an-enterprise-scale-agentic-ai-operating-system ; Wipro/Google Cloud Consulting video interview (90 agents / 15,000 associates, we are our own reference) https://www.technology-in-business.net/new-way-now-wipro-scales-agentic-ai-for-15000-associates-with-google-cloud-consulting/
FAQ
How many agents does McKinsey run internally?+
Per CEO Bob Sternfels on HBR’s IdeaCast: roughly 25,000 personalized agents beside 40,000 human employees, targeting near 1:1 by year-end; 1.5 million hours saved in search and synthesis last year; 2.5 million charts generated in six months — a figure not independently verified.
How large is EY’s agent deployment?+
Per EY’s official case study: more than 50,000 agents built in nine months; EY.ai EYQ deployed to 300,000+ professionals; the platform was co-built with Microsoft (orchestration) and Nvidia; EY says it drinks its own champagne before sharing lessons with clients.
What does Wipro mean by we are our own reference?+
Wipro CIO Kenny Kesar, in a Google Cloud Consulting video interview: an $11 billion services firm should incubate and validate solutions internally before selling them — we are our own reference. Wipro has deployed 90 orchestrated agents across 15,000 associates on Gemini Enterprise.
Why are these numbers called self-reported?+
McKinsey’s 25,000 agents, 1.5 million hours, and 2.5 million charts come from the firm’s own disclosures (the chart figure explicitly unverified); EY’s and Wipro’s figures come from vendor/partner case narratives with no third-party audit.
What does McKinsey’s own survey say?+
McKinsey’s 2026 State of AI global survey (1,719 respondents, May 4–June 8, 2026): 88% of organizations regularly use AI in at least one function, but only 39% attribute any EBIT impact (most below 5%), and only 6% qualify as AI high performers — flat versus 2025.
Which five questions should buyers ask an AI consultancy?+
① How many agents do you run internally, in which functions, for how long? ② Are they copilot-class or end-to-end? ③ How many hours did you save, on what baseline? ④ What are your error rates and escalation paths? ⑤ Are the internal numbers third-party audited? Specific answers with baselines and openness to verification win.
Related Articles
Agent ROI Is a Workflow-Design Problem, Not a Speed Race
First movers aren’t fastest to agent ROI; perfect data isn’t required; heavier governance slows ROI and doubles late error detection. Rollout sequence.
Agent scaling fails on allocation, not models
March 2026 survey (650 leaders): 78% pilot, 14% scale; 89% of failures trace to 5 gaps. Scalers fund eval and monitoring, not prompts. Scaling is allocation.
FDE arms race: 1,000 engineers in, who leaves the evidence?
Google rents Accenture for 1,000 FDEs; five armies compete on headcount. None answers who verifies agents post-handoff. Headcount is capacity, not proof.
Adoption sprints, EBIT stalls: a scaling-ROI gap
McKinsey 2026: agent scaling jumped 27% to 40%; EBIT impact stuck at 37%. High performers rebuilt workflows (73% vs 25%) — conversion, not tech.