Expand: New Source Leads, Grouped by the Gap They Close
Project Plan section 3.4 named three evidence gaps: no consistent definition or measure of AI value, little evidence directly linking trust and change readiness to financial results, and an evidence base dominated by Western consultancy research. The groups below attack those gaps, plus two more needs that surfaced in scoping: governance benchmarks for the maturity model, and fresher failure data. Each card states why it earns a place now.
Gap 1 · Causal, measured evidence of AI value (not self reported surveys)
The strongest weakness in the current base: almost every value number is a self reported executive survey. These are randomized and natural field experiments with measured outputs, the highest quality evidence available on what AI actually does to productivity and quality.
Navigating the Jagged Technological Frontier
Peer Reviewed JournalDell'Acqua, McFowland, Mollick, Lifshitz, Kellogg and colleagues ran a field experiment with 758 BCG consultants. Inside AI's capability frontier, consultants with AI finished more tasks, faster, at measurably higher quality; outside the frontier, AI use made results worse. Originally the famous 2023 HBS working paper, now published in a top peer reviewed journal.
Generative AI at Work
Peer Reviewed JournalBrynjolfsson, Li, and Raymond studied 5,172 customer support agents in a staggered real world rollout. AI assistance raised productivity 15% on average, with the largest gains for less experienced workers, improved customer sentiment, and improved employee retention.
Experimental Evidence on the Productivity Effects of Generative AI
Peer Reviewed JournalNoy and Zhang's randomized experiment on professional writing tasks: ChatGPT cut time taken by about 40% while raising output quality, and it compressed the gap between weaker and stronger performers rather than widening it.
The Cybernetic Teammate
Academic Working PaperDell'Acqua, Sadun, Lakhani and colleagues' follow up field experiment at Procter and Gamble: individuals working with AI matched the performance of two person teams without AI, and AI use broke down the usual silo between commercial and technical staff.
Gap 2 · Linking trust, readiness, and change to outcomes
The plan concedes that trust evidence measures perception, not results. These sources get closer to the behavior and leadership side of the link.
Superagency in the Workplace
Consultancy ResearchSurvey finding that employees are far more ready for AI than leaders believe: employees are already using AI for a large share of their work while leaders underestimate usage by roughly three times, and the biggest barrier to scaling is leadership, not employee resistance.
Work Trend Index
Vendor ResearchMicrosoft and LinkedIn's annual study of tens of thousands of workers across 31 countries, tracking real usage telemetry alongside survey data: employee adoption running ahead of official programs, the rise of human agent teams, and the emergence of frontier firms.
Gap 3 · Widening the lens beyond Western consultancy data
The plan names the Western, consultancy heavy skew as a limitation. These intergovernmental datasets cover most of the world with published methods and no product to sell.
AI Preparedness Index
Intergovernmental DatasetIMF index covering 174 economies across digital infrastructure, human capital and labor policies, innovation, and regulation. Shows advanced economies better positioned but more exposed to disruption, while emerging markets face a widening readiness gap.
OECD.AI Policy Observatory
Intergovernmental ObservatoryLive database of national AI strategies, policies, and incidents across 70 plus jurisdictions, plus the OECD AI Principles that most national frameworks trace back to. Includes trend data on investment, skills, and regulation by country.
Gap 4 · Governance benchmarks the maturity model can anchor to
Dr. Kokkonen asked for a maturity model with stage gates. These three are the recognized external benchmarks a governance dimension can be graded against, which keeps the framework from being arbitrary.
AI Risk Management Framework (AI RMF)
Government FrameworkThe US national framework for trustworthy AI: Govern, Map, Measure, Manage. Voluntary, widely adopted by enterprises as the de facto governance baseline, with a Generative AI profile added for GenAI specific risks.
ISO/IEC 42001:2023 AI Management Systems
International StandardThe first certifiable international standard for AI management systems, covering governance, risk, lifecycle management, and continual improvement. Organizations can now be audited and certified against it.
The EU AI Act: Regulatory Framework on AI
RegulationThe world's first comprehensive AI law, phasing in through 2026 and 2027 with risk tiered obligations and significant penalties. The reference point for Europe's regulation first posture and for compliance as a value protection activity.
Gap 5 · Fresher failure and abandonment data
The initial base leans on the contested MIT 95% figure and Gartner's forward prediction. These add measured abandonment rates and systematically studied root causes.
The Root Causes of Failure for AI Projects
Policy Research InstituteRAND interviewed 65 experienced data scientists and engineers and synthesized five root causes of AI project failure, led by leaders misunderstanding or miscommunicating the problem to be solved, with inadequate data and infrastructure close behind. Technology shortcomings rank last.
AI Abandonment Jumps from 17% to 42%
Media Summary of SurveyS&P's survey of 1,000 plus enterprises across North America and Europe: the share of companies abandoning most of their AI initiatives jumped from 17% in 2024 to 42% in 2025, with an average 46% of proofs of concept scrapped before production. Cost, data privacy, and security cited as top causes.
30% of GenAI Projects Abandoned After Proof of Concept
Analyst ResearchThe predecessor to the agentic prediction already carded: Gartner forecast at least 30% of generative AI projects would be abandoned after proof of concept by end of 2025, citing poor data quality, inadequate risk controls, escalating costs, and unclear business value.
Primary evidence the team can create
Expanding is not just finding documents. The team's own position gives it primary evidence no consultancy has, and Ducere explicitly invites surveys and personas as deliverables. Anything involving people goes through the plan's section 5 ethics and approval gate first.
Practitioner Interview Program
Primary EvidenceSix to ten short semi structured interviews with leaders and practitioners across the team's networks, using the LinkedIn channel Dr. Kokkonen opened (she accepts all connection invites and encouraged cross group networking). Target a spread of industries, regions, and seniority, asking how they define AI value, where it leaked, and what they would measure differently.
Stakeholder Perception Pulse Survey
Primary EvidenceA short anonymous survey built on the Melbourne and KPMG question structure, distributed through the team's combined networks across the six RFS stakeholder groups. Even 50 to 100 responses gives the perception heat map an original data layer to compare against the global benchmarks.
DXC Public Materials Sweep
Primary EvidenceA systematic pass through DXC's own public record: the full AdvisoryX study download, annual report and investor materials (how DXC describes AI demand to shareholders), newsroom releases, and Dr. Kokkonen's publication and speaking trail including her Trust in Government's use of AI talk.
Cross Industry Mini Cases from the Team Itself
Primary EvidenceFive members spanning software sales (UK and global), digital marketing operations (London), government contracting operations (Hawaii), and construction and refrigeration services (Dallas). Each member writes a one page structured vignette of AI value creation or value loss observed in their own sector, using a common template.
Assess: The Quality Screen Every Source Passes Before It Enters the Base
Run each new source through the six checks below before it is added to a tab. A source can still be used if it fails a check, but the weakness gets written down next to it, the way section 3.1 of the Project Plan already does with its Critical Consideration column. The standing rule from section 3.4: consultancy and vendor claims are triangulated against academic sources before they carry weight in findings.
| Check | Question to Ask | Passes When | Red Flags |
|---|---|---|---|
| Credibility | Who produced it and what is their track record? | Named authors or established institution; peer review or published editorial standards. | Anonymous content, AI generated blog farms, no institutional home. |
| Currency | Is it current enough for a field moving this fast? | 2024 or later for market data; older is fine for theory (J curve, trust theory) where the concept, not the number, is being used. | Pre 2023 adoption or ROI statistics presented as current. |
| Relevance | Which research question or objective does it serve? | Maps cleanly to at least one of Q1 to Q4, the international lens, or a maturity model dimension. | Interesting but unmappable; collect somewhere else, not in the base. |
| Method transparency | Can you tell how the numbers were produced? | Sample size, population, and method published; limitations acknowledged. | Headline percentages with no methodology anywhere (the DXC AdvisoryX study itself carries this caution). |
| Commercial interest | Does the publisher sell a remedy for the problem it describes? | No stake, or the stake is declared and the claim is triangulated against an independent source. | Vendor research used alone to justify a finding that benefits the vendor. |
| Primary vs secondary | Is this the original source or a retelling? | Original report, paper, or dataset cited; media retellings used only for critique and color. | Citing a news summary when the underlying report is freely available. |
Organize: Evidence Coverage Against the Project's Needs
Objectives O1 to O5 are still being finalized in section 1.5 of the Project Plan, so coverage is organized against the four RFS research questions plus the three framework needs the sponsor set. When the objectives are locked, this table gets recut against them. Status reflects the base as of 3 October: the job of this phase is to move every row to Strong by 18 October.
| Project Need | Status | What Is Already Strong | Gap to Close by 18 Oct | New Leads on This Page |
|---|---|---|---|---|
| Q1 What is AI value | Adequate | Survey evidence (BCG, McKinsey, Deloitte) and the academic definition (Enholm). | Causal, measured evidence; a working definition tested against it. | Jagged Frontier, Generative AI at Work, Noy and Zhang, Cybernetic Teammate. |
| Q2 What AI is not | Strong | Gartner agent washing, AI theater, hype critique both directions. | Academic grounding for capability limits. | Jagged frontier concept (capability is uneven, not general). |
| Q3 How AI is perceived | Adequate | Melbourne and KPMG global data, Stanford perception gap, trust theory (Glikson and Woolley). | Perception linked to behavior and outcomes; original stakeholder data. | Superagency, Work Trend Index, pulse survey, practitioner interviews. |
| Q4 Why businesses struggle | Strong | DXC five barriers, MIT learning gap, workflow redesign evidence, CEO testimony. | Non commercial confirmation; measured (not predicted) failure rates. | RAND root causes, S&P 42% abandonment, Gartner 2024 baseline. |
| Value over time horizons | Adequate | J curve theory, Deloitte one to five year agentic expectations. | Firm level evidence of the lag and recovery. | Field experiments above give the early, task level end of the curve. |
| International perspective | Needs More | Adoption and trust percentages by country. | Non Western, non consultancy data; regulation as a regional cost factor. | IMF Preparedness Index, OECD.AI Observatory, EU AI Act. |
| Maturity model and diagnostic inputs | Needs More | DXC five barriers and Enholm enablers as candidate dimensions. | External benchmarks so maturity grades are defensible, not invented. | NIST AI RMF, ISO/IEC 42001, EU AI Act risk tiers, Edge Framework phases. |
New References Added This Phase (Harvard Style)
Only the sources introduced on this tab. The foundation references stay on the Initial Research tab. Verify dates and details against each original publication before final submission.
Brynjolfsson, E., Li, D. and Raymond, L. (2025) 'Generative AI at Work', The Quarterly Journal of Economics, 140(2), pp. 889 to 942. doi: 10.1093/qje/qjae044.
Dell'Acqua, F., McFowland, E., Mollick, E., Lifshitz, H., Kellogg, K.C., Rajendran, S., Krayer, L., Candelon, F. and Lakhani, K.R. (2026) 'Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality', Organization Science. doi: 10.1287/orsc.2025.21838.
European Commission (2024) Regulatory Framework on Artificial Intelligence. Available at: https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai (Accessed: 3 October 2026).
Gartner (2024) Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept by End of 2025 [Press release]. Available at: https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025 (Accessed: 3 October 2026).
Harvard Business School AI Institute (2025) The Cybernetic Teammate: A Field Experiment on Generative AI Reshaping Teamwork and Expertise. Available at: https://aiinstitute.hbs.edu/?p=26305 (Accessed: 3 October 2026).
International Monetary Fund (2024) AI Preparedness Index. Available at: https://www.imf.org/external/datamapper/datasets/AIPI (Accessed: 3 October 2026).
International Organization for Standardization (2023) ISO/IEC 42001:2023 Information Technology, Artificial Intelligence, Management System. Available at: https://www.iso.org/standard/81230.html (Accessed: 3 October 2026).
McKinsey & Company (2025) Superagency in the Workplace: Empowering People to Unlock AI's Full Potential. Available at: https://www.mckinsey.com/capabilities/mckinsey-digital/our-insights/superagency-in-the-workplace-empowering-people-to-unlock-ais-full-potential-at-work (Accessed: 3 October 2026).
Microsoft (2025) Work Trend Index Annual Report. Available at: https://www.microsoft.com/en-us/worklab/work-trend-index (Accessed: 3 October 2026).
National Institute of Standards and Technology (2023) AI Risk Management Framework (AI RMF 1.0). Available at: https://www.nist.gov/itl/ai-risk-management-framework (Accessed: 3 October 2026).
Noy, S. and Zhang, W. (2023) 'Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence', Science, 381(6654), pp. 187 to 192. doi: 10.1126/science.adh2586.
OECD (2026) OECD.AI Policy Observatory. Available at: https://oecd.ai (Accessed: 3 October 2026).
RAND Corporation (2024) The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed. Available at: https://www.rand.org/pubs/research_reports/RRA2680-1.html (Accessed: 3 October 2026).
This Week Health (2025) AI Project Failures Surge to 42% as Companies Struggle to Scale [Summary of S&P Global Market Intelligence survey]. Available at: https://thisweekhealth.com/news_story/ai-project-failures-surge-to-42-as-companies-struggle-to-scale (Accessed: 3 October 2026).