FY2027 is the first fiscal year in which most CFOs will treat AI as a line item to be managed rather than a bet to be defended — and this playbook is about doing that without strangling the programs that actually work.
Why FY2027 AI Budgeting Is a Different Exercise
For two budget cycles, AI spending lived in a protected category. Boards understood it as strategic exploration; CFOs signed off on pilots the way they sign off on R&D — knowing most attempts fail, but that the option value justifies the cost. That era is ending, and the numbers explain why. Gartner (2025) estimates over 40% of agentic AI projects will be canceled by 2027; a widely cited MIT study (2025) put the share of generative AI pilots with no measurable P&L impact near 95%. When failure rates like that become public knowledge, "trust the exploration" stops being a viable budget posture.
FY2027 budgeting is therefore different in kind, not just degree. Three structural changes define it:
- AI moves from project funding to run-rate visibility. What used to be a one-time pilot budget now includes recurring model, infrastructure, and platform costs that behave like software subscriptions — except they scale with usage, not seats.
- Finance gets a seat at the architecture table. Decisions such as which models to route to, whether to self-host, and how to meter internal consumption now directly determine cost curves. These are finance questions as much as technology questions.
- The unit of analysis shifts from "AI initiative" to "decision improved." The FY2027 test for funding is no longer "is this innovative?" but "does this measurably make a specific decision faster, cheaper, or better — and can we prove it with a number we agreed on in advance?"
McKinsey (2025) found that only about a quarter of AI-adopting organizations report enterprise-level EBIT impact. The implication for budget planning is uncomfortable but clarifying: most of what your organization currently spends on AI is, by the evidence, not yet producing returns. The playbook below is about finding which part is, funding it properly, and stopping the rest.
The Structural Split: Platform Versus Pilots
The first architectural decision of an FY2027 AI budget is not a technology decision at all — it is how much of the budget goes to experimentation versus how much goes to industrialization. Most FY2025 and FY2026 AI budgets were structurally skewed: 70–90% of spend sat in pilots, each with its own tooling, its own data pipeline, and its own evaluation standard. That structure produced exactly what it was designed to produce — many experiments and few production systems.
For FY2027, the healthier structure inverts the ratio for organizations past the discovery phase. A practical reference frame:
| Budget component | Typical FY2026 share | Recommended FY2027 share | What it funds |
|---|---|---|---|
| Experiments and new pilots | 60–80% | 15–25% | New use-case validation with fixed-price, time-boxed trials |
| Production platform and data readiness | 10–20% | 35–45% | Semantic layers, permissions, connectors, integration (e.g., MCP-based), observability |
| Operations and governance | 5–10% | 20–30% | Monitoring, evaluation harnesses, audit trails, model cost management |
| Skills and change adoption | 5–10% | 10–15% | Training decision-makers and analysts on AI-delivered workflows |
The logic behind the shift: platforms amortize across use cases, pilots do not. Every pilot that hard-codes its own data integration is a tax on every future pilot. The platform share of the budget is what makes the tenth use case cost 20% of what the first one did — and that cost curve, not any individual use case, is where durable value concentrates.
The failure mode to avoid is the mirror image: over-consolidating before any pilot has proven value. If your organization has not yet demonstrated a single AI use case with defensible unit economics, do not fund a platform at 40% — fund two or three falsifiable pilots and let the evidence set the allocation. The split above assumes discovery is done; if it is not, finish it cheaply first.
New Line Items: Modeling LLM and Token Operating Costs
The FY2027 budget contains cost lines that had no equivalent two years ago, and most organizations are still modeling them badly. Token-based model costs behave unlike any legacy IT cost: they scale with conversational volume, are highly sensitive to architecture choices, and can swing 10x based on decisions made before the first invoice arrives.
Break the model cost line into four components:
- Inference spend per use case. Not one "AI cost" line — one line per production use case, metered. A customer-service summarization use case and an analyst copilot have wildly different token profiles and must be budgeted separately.
- Model tiering. Route simple queries to cheaper models and reserve premium models for tasks that need them. In practice, well-architected deployments route 60–80% of requests to small or mid-tier models; the cost delta between routing everything to a frontier model and tiering intelligently is routinely 3–5x on the same workload.
- Context and retrieval costs. Every query that drags a large context window through the model pays for it. Optimizing retrieval — sending less, more relevant data — is a cost-control lever, not just a quality lever.
- Evaluation and guardrail costs. Running evals, red-teaming, and monitoring generates its own inference spend. It is small relative to production traffic but must be named, or it gets silently absorbed and resented.
Two budgeting rules keep this manageable. First, require every AI line item to have a usage cap and a budget alert at 80% — surprises, not absolute levels, are what damage finance's relationship with AI programs. Second, negotiate committed-use pricing once a use case's monthly volume is stable; frontier-model pricing has fallen sharply through 2024–2026, and organizations that renegotiated mid-cycle captured substantial savings. IDC's (2024) spending forecasts, which show platform and services spend growing faster than raw model APIs, reflect exactly this maturation.
Unit Economics: The CFO's Core AI Metric
Every FY2027 AI funding decision should reduce to one question: what does one unit of this workload cost, and what is that unit worth? The discipline is borrowed from SaaS and marketplace businesses, and it transfers cleanly.
Define the unit per use case. For a conversational BI deployment, the unit might be "one data question answered with a permissioned, sourced answer." For a document-processing pipeline, "one contract extracted and validated." For a service copilot, "one customer interaction assisted." Then build the simple table finance can actually govern:
| Use case | Unit | Unit cost (modeled) | Baseline cost of the same unit, manual | Monthly volume | Monthly net benefit at steady state |
|---|---|---|---|---|---|
| Conversational BI (IM-native) | Answered data question | HKD 0.4–1.2 | HKD 15–30 (analyst time) | 40,000 | Meaningful, staff-time led |
| Contract extraction | Document processed | HKD 2–5 | HKD 60–120 (reviewer time) | 3,000 | Meaningful, cycle-time led |
| Service interaction copilot | Assisted interaction | HKD 0.8–2.0 | HKD 8–15 (handle-time delta) | 25,000 | Meaningful, quality + handle-time led |
The numbers above are illustrative modeling anchors, not benchmarks — but the shape is the point. Note two properties. First, unit costs are cents against baseline costs of dollars; the economics of AI use cases are almost always dominated by the human time they displace or accelerate, not by token prices. Second, volume drives everything: a use case with excellent unit economics at 40,000 monthly queries is irrelevant at 400, which is why adoption (not accuracy) is the most common reason AI programs miss their business case.
McKinsey's (2025) finding that only ~25% of AI adopters report enterprise-level EBIT impact is, in unit-economics terms, a volume problem more than a cost problem. The FY2027 budget should therefore fund adoption mechanisms — embedding AI where work already happens, for instance in messaging platforms — with the same seriousness as the technology itself.
Defunding Zombie Pilots: A Principled Cull
Every enterprise carrying more than five AI initiatives has zombies: pilots that finished their evaluation window months ago, produce no decision-relevant metric, and survive because nobody owns the decision to stop. Gartner (2024) predicted ~30% of generative AI projects would be abandoned after proof of concept; the FY2027 budget cycle is where your organization formally performs that abandonment, on purpose, rather than by attrition.
A principled cull has three steps. Step one: publish the criteria before reviewing anything. Suggested bars: each initiative must name its owner, the decision it improves, its monthly run-rate cost, and one pre-agreed measurable result. Anything missing two of the four enters the review queue. Step two: review against those criteria in one sitting. The review is administrative, not political — the criteria do the deciding. Step three: sunset with a deadline, and reallocate visibly. Freed budget should flow to the platform and operations lines described above, in the same cycle, so the cull visibly funds something rather than just removing something.
Two failure modes deserve names. The "sunk-cost zombie" survives because its sponsor has reported on it for a year; the fix is requiring a falsifiable result, not progress narratives. The "stealth zombie" survives because its run-rate is hidden inside a cloud bill or a team's operating budget; the fix is the complete inventory — every AI initiative, one page, four fields. A widely cited MIT study (2025) put the share of valueless generative AI pilots near 95%; your organization's number will be lower, but it will not be zero, and the FY2027 budget should be built as if it will not be small.
Procurement Checkpoints: Buying AI Without Buying Surprises
AI procurement has features that classic software procurement handles poorly: usage-based pricing that can scale unpredictably, model behavior that changes with vendor updates, and data flows that cross organizational boundaries. Five checkpoints, written into the standard process for FY2027, close most of the gaps:
- Checkpoint 1 — Data flow disclosure. Where does data transit, where is it stored, is it used for vendor model training, and how is access revoked? Make the answer a contract annex, not a sales-deck assurance.
- Checkpoint 2 — Permission architecture. The system must enforce the requesting user's data permissions at query time. Any tool that answers from a god-mode service account is disqualified for anything touching regulated data — IBM (2025) found 97% of organizations with AI-related breaches lacked AI access controls.
- Checkpoint 3 — Priced usage model. Vendor pricing must be expressible as cost per unit of the workload (per query, per document), with volume bands and caps. If a vendor cannot state unit economics, the buyer will own the surprise.
- Checkpoint 4 — Evaluation rights. Contractually reserve the right to run pre-agreed accuracy tests against your own data before renewal. Model updates can silently change behavior; the renewal decision needs evidence.
- Checkpoint 5 — Exit and portability. Define what happens to conversation logs, embeddings, and configuration on termination. This is cheap to negotiate at signature and nearly impossible to renegotiate later.
For time-boxed purchases — pilots especially — the structure finance should prefer is small, dated, and falsifiable. A fixed-price two-week pilot in the HKD 25k range that tests a specific workflow against pre-agreed metrics is a better financial instrument than a six-month exploratory engagement at ten times the cost, because the downside is bounded and the information yield per dollar is higher.
Value Gates: Making AI Spend Falsifiable
The last piece of the playbook is the mechanism that makes everything above enforceable: pre-agreed value gates attached to each tranche of AI spend. A value gate is a commitment written before deployment — if the metric misses the bar at the checkpoint date, the next tranche does not release. It converts AI budgeting from a narrative exercise into a controlled experiment, which is the only form finance can govern at scale.
Design rules for value gates:
- Metrics must be decision-level, not model-level. "95% answer accuracy" is a model metric; "store managers use it daily and reorder decisions happen 2 days earlier" is a decision metric. Fund the second.
- Baselines are measured, not asserted. Measure the current cost and cycle time of the target decision in the two weeks before deployment. A gate with an unmeasured baseline is a gate that can be renegotiated — which means it is not a gate.
- Checkpoints are short. Thirty to sixty days, aligned to how fast usage data accumulates. Annual value gates are where AI programs go to die quietly.
- The consequence is real. A gate that misses must release nothing — not "with exceptions," not "pending a plan." One enforced gate buys more credibility for the whole AI portfolio than ten gates that bend.
McKinsey (2025) and IDC (2024) both point the same direction: AI value is concentrating in organizations that industrialize measurement. The value-gate mechanism is how a CFO imports that discipline without needing to personally evaluate retrieval architectures — the gates do the governing.
Who Owns What: Governing the Budget Process Itself
The mechanisms above — the platform/pilot split, unit economics, value gates — fail quietly when nobody owns the process that enforces them. FY2027 is the year to name that owner and write down the operating rhythm, because AI spending now crosses enough organizational boundaries that leaving it unassigned guarantees drift.
A workable ownership model keeps roles narrow:
- The CFO owns the gates. Not the metrics themselves — business owners propose those — but the enforcement: whether a missed gate stops money. This is the one accountability that cannot be delegated downward without dissolving.
- The CIO or CDO owns the platform line. Data readiness, connectors, permissions, and the delivery surface (including IM-native channels) are infrastructure decisions with 3–5 year cost implications; they should not be re-litigated per use case.
- Business owners own unit economics. The head of supply chain, not IT, declares what a faster replenishment decision is worth. Finance validates the arithmetic; it does not invent the value.
- Internal audit owns the governance checklist. Data flow disclosure, access controls, and audit trails get reviewed by the function whose finding actually hurts — which is what makes the checklist credible rather than ceremonial.
The operating rhythm is light but non-negotiable: a monthly one-hour AI spend review covering unit costs against budget, active usage against targets, and gate status per use case; a quarterly reallocation decision where freed budget moves between lines. McKinsey's (2025) finding that only ~25% of AI adopters report enterprise-level EBIT impact is, in many cases, an ownership failure before it is a technology failure — programs with no named owner drift toward exactly the unmeasured spend the 75% exhibit.
One caution: do not create a new committee for this. The monthly review should live inside the existing IT or transformation steering forum, with AI as a standing agenda item. New governance bodies multiply meetings without multiplying decisions, and FY2027 budgets fund decisions.
An Illustrative FY2027 Allocation
The table below sketches a reference allocation for a mid-size enterprise (roughly HKD 12–18M annual AI budget) that has completed discovery and has 2–3 use cases ready to industrialize. It is a starting point for negotiation, not a prescription.
| Budget line | Share | Example items | Governing value gate |
|---|---|---|---|
| Production platform & data readiness | 40% | Semantic layer, permission-aware connectors, IM-native delivery (WeChat Work/DingTalk/Feishu/WhatsApp/Teams), observability | Tenth use case deploys at ≤30% of first use case's marginal cost |
| Industrialized use cases (run) | 25% | 2–3 production workloads: inference, integration, support | Unit cost per workload within ±20% of modeled budget, monthly |
| Governance, security, evaluation | 15% | Audit trails, access controls, eval harnesses, monitoring | Zero critical audit findings; evals run monthly per use case |
| New experiments | 15% | Fixed-price, time-boxed pilots (e.g., 2-week HKD 25k format) | Each pilot passes or fails a pre-agreed 30-day metric |
| Skills & adoption | 5% | Training, workflow redesign, champion networks | Active-usage rate ≥60% of licensed users by day 60 |
Three notes on reading the table. First, the experiment line is capped by design — this is the structural protection against the pilot sprawl that consumed FY2025 and FY2026 budgets. Second, the adoption line is small in money but outsized in leverage; McKinsey's (2025) 25% figure is substantially an adoption failure, and five percent of budget spent on workflow redesign is often the highest-return line in the table. Third, every line has a gate — because in FY2027, the CFO's job is not to decide whether the organization believes in AI, but to ensure every remaining dollar of AI spend is still earning its place.