AI Workflow Agency
AI5 min read

Enterprise AI Consultancy: How to Choose, Scope and Get Value

A practical guide to enterprise AI consultancy: what good looks like, how to scope engagements, pricing benchmarks, procurement and delivery models

By AI Advisory team

Most enterprise AI engagements still end as a slide deck. McKinsey's 2024 State of AI survey found that 72% of organisations have adopted AI in at least one function, yet only around 8% of respondents say enterprise-wide AI efforts have driven meaningful EBIT contribution. The gap sits between advisory work that stops at strategy and delivery work that starts without one. This guide covers what an enterprise AI consultancy actually does, how to scope engagements so they ship, what to pay, and how to avoid the failure modes that eat most first-year budgets.

What an enterprise AI consultancy actually does

The term covers a wide range of firms, from Big Four advisory practices running multi-million-pound transformation programmes down to specialist boutiques building a single RAG pipeline. Useful distinctions:

  • Strategy-only firms produce readiness assessments, opportunity maps, target operating models and business cases. They rarely ship code. Deloitte, BCG, McKinsey, PA Consulting and Bain sit here for most engagements.
  • Systems integrators deliver at scale using vendor platforms (Microsoft, Google Cloud, AWS, Databricks, Salesforce). Accenture, Capgemini, Infosys and TCS dominate. Strong on delivery capacity, expensive on day rates, cautious on emerging tooling.
  • Build-focused specialists combine a smaller strategy footprint with hands-on engineering. They own the roadmap and the codebase. Team sizes are 10-80, engagements £15k-£500k, delivery cycles 8-16 weeks.
  • Product engineering firms take a fixed problem and ship. Weaker on discovery, strong on velocity. Best used after the strategy is settled.

The right choice depends on where the risk sits. If the business case is unclear, buy strategy. If the case is clear but the internal team cannot ship, buy build. If both are unclear, buy a firm that does both under one roof, and insist on a working prototype inside the first eight weeks.

The four engagement types worth knowing

Enterprise AI work tends to arrive in one of four shapes. Getting the shape right at the RFP stage saves months of misalignment.

1. Readiness and opportunity assessment

Two to six weeks. Fixed fee, typically £15k-£60k. Outputs are a prioritised opportunity list scored on value, feasibility and time-to-value, a data and tooling audit, a target architecture sketch, and a costed 12-month roadmap. This is the right first engagement when leadership disagrees on where to start, or when the board has asked for an "AI strategy" without a specific use case in mind.

2. Proof of value

Four to twelve weeks. Fixed scope, typically £30k-£120k. One use case, one working system, real users, measurable output. Not a demo. The deliverable is running software with a plan to move to production if the metrics land. Around 30% of proofs of value stall at this gate, which is fine - the point is to fail cheaply before committing to the larger build.

3. Production build

Eight to twenty-six weeks. Typically £80k-£500k. Full delivery: infrastructure, integrations, evaluation harness, security review, go-live, handover. Retainer usually follows for operation and iteration.

4. Managed operation

Ongoing. £5k-£40k per month depending on system complexity and SLAs. Model monitoring, prompt and pipeline iteration, incident response, cost optimisation, quarterly roadmap review. Around 70% of build engagements convert to retainer in our own book, and the pattern is common across the mid-market segment.

Where enterprise AI actually pays back

Ignore the marketing decks. The use cases that consistently deliver measurable ROI in mid-market and enterprise deployments cluster in a small number of areas.

Document-heavy operations. Contract review, claims processing, KYC, supplier onboarding, RFP response drafting. A well-built RAG assistant on a corpus of 10,000+ documents typically cuts handling time by 40-70%. The Bank of England's 2024 survey of AI use in UK financial services found risk and compliance functions among the most common deployment areas, precisely because the documents are structured enough to retrieve against.

Customer service triage and deflection. Not full replacement - triage. A retrieval-grounded assistant that resolves tier-1 queries and hands off cleanly to humans on tier-2 typically deflects 25-45% of ticket volume within six months of deployment. Klarna's widely-cited assistant handled work equivalent to 700 full-time agents within its first month, though the company later hired back human agents for higher-complexity work.

Sales operations. Lead scoring, CRM enrichment, meeting summarisation, proposal drafting. Payback is fast because the baseline is manual work by expensive people. HubSpot and Salesforce integrations are the usual delivery surface.

Internal knowledge search. The classic "where is that policy document" problem. Cheap to build, universally useful, and a good first project because the failure cost is low.

Workflow automation with AI in the loop. The vast majority of enterprise value still comes from removing manual data movement between systems. Tools like n8n, Make and Zapier handle the plumbing; LLMs handle the judgement calls (classification, extraction, drafting) inside the flow.

Use cases to be sceptical of: fully autonomous agents making irreversible decisions, marketing content generation at scale without editorial oversight, and anything pitched as "replacing" a function rather than augmenting it.

What to pay, and how

Day rates in the UK enterprise AI consultancy market as of 2025:

  • Big Four and MBB strategy work: £2,500-£4,500 per consultant per day, with partner rates above £6,000.
  • Tier-one systems integrators: £1,200-£2,200 per day blended, £2,500+ for senior architects.
  • Specialist AI boutiques: £900-£1,600 per day for senior engineers, £1,400-£2,200 for principals.
  • Offshore or nearshore delivery: £350-£700 per day. Quality varies enormously; the cheap end is a false economy for mid-market work.

Rate cards matter less than delivery model. The three commercial structures worth considering:

Fixed price, fixed scope. Best for readiness assessments and proofs of value where the deliverable is knowable. Removes surprise, but change requests are painful. Insist on a written change control process.

Time and materials with a cap. Best for production builds where requirements will evolve. The cap protects you; the T&M keeps the vendor honest on scope creep. Expect weekly burn-down reporting.

Outcome-based pricing. Rare, and mostly theatre. Genuine outcome pricing requires a baseline both parties trust and a clean attribution model. If a vendor offers it without discussing baseline measurement in detail, they are proposing T&M with a marketing wrapper.

For a first engagement, our default recommendation is: fixed-fee discovery (£20k-£40k), fixed-fee proof of value (£40k-£100k), then T&M-with-cap production build. Do not sign a six-figure production contract before you have working software from the same team.

Procurement, security and GDPR

Enterprise procurement is where good projects go to die. A few things worth getting right early:

Data processing agreement (DPA). Any consultancy touching personal data needs a DPA with clear sub-processor disclosure. If they route data through OpenAI, Anthropic, Google or Azure OpenAI, that must be named. The ICO's guidance on AI and data protection is the reference document for UK deployments.

Model and data residency. For regulated sectors, insist on EU or UK residency for both training data and inference. Azure OpenAI Service in UK South and Anthropic's EU-hosted endpoints via AWS Bedrock are the usual answers. OpenAI's Enterprise plan offers zero-retention and EU data residency as of 2024.

DPIA. A Data Protection Impact Assessment is required under UK GDPR for most systematic AI processing of personal data. Your consultancy should offer to run one, or work with your DPO to complete it. If they have not heard of a DPIA, walk away.

Cyber Essentials Plus, ISO 27001, SOC 2. Standard asks for any vendor handling production data. Cyber Essentials Plus is the minimum for UK public-sector work.

IP and model ownership. Get the contract to say clearly who owns prompts, fine-tuned model weights, evaluation datasets and code. Default should be client-owned with a licence-back to the consultancy for reference-architecture reuse.

Exit clauses. A production system built on someone else's infrastructure is a lock-in risk. Insist on portable deployment (containerised, documented, deployable to your own cloud account), source code escrow if the consultancy is small, and a defined handover package.

How to run the first 90 days

The first quarter determines whether the engagement produces value or joins the 70% of enterprise AI projects that BCG's 2024 Build for the Future research found deliver disappointing results.

Week 1-2: Discovery. The consultancy interviews 8-15 stakeholders, audits the data and tooling estate, and produces a ranked opportunity list. You review and pick one use case for the proof of value. Do not pick more than one.

Week 3-4: Specification. Written specification covering user journey, success metrics, data sources, integrations, evaluation approach, and go/no-go criteria for production. Signed by the business owner, not just IT.

Week 5-10: Build. Working software demonstrable by week six. Weekly demos, not weekly status reports. If you are on week seven with no software to touch, escalate.

Week 11-12: Evaluation. Real users, measured against the success metrics defined in week four. Go/no-go decision on production.

Week 13+: Production or pivot. Either move to production build with the same team, or take the learnings and pick a different use case. Do not extend a failed proof of value in hope.

The most common failure mode is not technical. It is that the business owner disappears after kickoff and the consultancy is left interpreting requirements from middle management. Get the executive sponsor on the demo calls from week one.

Signals of a consultancy worth hiring

Beyond the obvious (case studies, references, security certifications), a few less-obvious signals:

  • They ask about your data before pitching a use case. If the first meeting is a slide deck on generative AI trends, you are being sold to, not consulted.
  • They name specific tools and versions. "We use n8n self-hosted on Hetzner with Postgres for state, pgvector for retrieval, and Claude Sonnet for extraction" is a real answer. "We are technology-agnostic" is a hedge.
  • They will show you code. A build-focused consultancy should be happy to walk through a previous client's repository (redacted) on the first technical call.
  • They quote realistic timelines. Anyone promising a production RAG system in three weeks has not built one that survived contact with real users.
  • They talk about evaluation. Systems without evaluation harnesses drift silently. If the proposal has no line item for eval, the vendor plans to hope.
  • They can articulate what they will not do. A consultancy that says yes to everything will do most things badly.

Frequently asked questions

How is an enterprise AI consultancy different from a management consultancy?

Management consultancies (McKinsey, BCG, Bain, Big Four) focus on strategy, operating model design, and change management. Their deliverable is typically a set of recommendations and a business case. An enterprise AI consultancy in the modern sense combines that advisory work with engineering delivery - the same team that scopes the roadmap also builds the software. This matters because AI strategy done in isolation from build reality tends to over-promise; the specific limitations of retrieval quality, latency, model cost and data readiness only surface once you try to ship something. Firms that do both under one roof avoid the handover gap where 60-70% of enterprise AI initiatives stall.

What should a first engagement cost?

For a mid-market business (50-1000 employees), a sensible first engagement is a two-to-four-week readiness assessment at £15k-£40k, followed by a proof of value at £40k-£120k if the assessment identifies a strong candidate use case. Avoid signing a production build contract above £100k before you have working software from the same team on a smaller scope. Big Four firms will quote 3-5x these numbers for the same deliverables; specialist boutiques quote 30-50% less than tier-one systems integrators for equivalent work. The right benchmark is not day rate but total cost to first production deployment.

Should we build in-house or hire a consultancy?

Build in-house when you have three or more senior engineers who can commit 50%+ of their time for two quarters, an existing MLOps or platform team to operate what gets built, and a specific use case with clear requirements. Hire a consultancy when you need to move faster than hiring allows, when the use case is exploratory (you might build the wrong thing), or when you need cross-domain expertise (retrieval + evaluation + integrations + security) that is expensive to assemble in one internal team. Many mid-market clients run a hybrid: consultancy builds v1 and trains the internal team, who then operate and iterate.

How do we protect ourselves from vendor lock-in?

Insist on four things in the contract. First, code and configuration are client-owned and delivered to a client-controlled repository from day one. Second, the deployment target is your cloud account (AWS, Azure, GCP) or a documented on-premise setup, not the vendor's infrastructure. Third, the consultancy provides a written handover package covering architecture, runbooks, credentials rotation, and incident response. Fourth, there is a defined transition-out clause with a fixed-fee cap for handover to a successor supplier. If a consultancy resists any of these, they are planning to make themselves hard to replace, which is a red flag regardless of quality.

How do we measure whether the engagement is working?

Define success metrics before build starts, not after. For a customer service assistant, that is deflection rate, CSAT on AI-handled conversations, and cost per resolved ticket. For a document processing system, it is handling time, accuracy against a golden dataset, and exceptions per thousand documents. For sales enablement, it is meeting-to-opportunity rate and rep-hours saved per week. Weekly demos with real usage data, not slide-based status reports, are the operational indicator. If you cannot see users interacting with working software by week six of a proof of value, something is wrong. If the vendor cannot articulate the eval methodology, that is worse.

What about GDPR and the EU AI Act?

UK GDPR still governs personal data processing in the UK, and the ICO's guidance on AI is the reference document. Your consultancy should complete a DPIA for any systematic processing of personal data, name all sub-processors (including LLM providers), and use UK or EU data residency where sensitivity warrants. The EU AI Act, in force from August 2024 with phased application through 2026-2027, applies if you place systems on the EU market or process EU residents' data. Most mid-market use cases fall into the "limited risk" or "minimal risk" tiers with transparency obligations only. High-risk applications (credit scoring, employment decisions, critical infrastructure) trigger substantially heavier compliance burden and should be scoped with legal counsel from week one.

Can a small consultancy really deliver enterprise-grade work?

Yes, with two caveats. First, check that they have shipped comparable systems to production, not just prototypes - ask for references at organisations of similar size and complexity to yours. Second, check the operational side: incident response, on-call coverage, escalation paths, security certifications. A ten-person consultancy can build a better RAG system than a thousand-person integrator, but a ten-person consultancy cannot offer 24/7 sub-hour incident response without partners. Match the operational model to your risk tolerance. For a customer-facing system with revenue implications, that means either a larger vendor or a specialist boutique with a named managed-service partner.

What if we have no data strategy yet?

Start there. AI systems are downstream of data quality, and a consultancy that agrees to build on top of a broken data foundation is either optimistic or dishonest. A useful sequence for organisations without a data strategy: three-week data audit (what exists, where, in what state, who owns it), followed by a targeted cleanup on the two or three sources needed for the first AI use case, followed by the AI build itself. This is not the exciting sequence, but it is the one that ships. Attempting to run an enterprise-wide data cleanup before any AI work is delivered is the other common failure mode - three years and no output.

Closing thought

Enterprise AI consultancy done well looks unglamorous: fixed-scope discovery, one use case at a time, working software by week six, evaluation harnesses that catch drift, and a handover package that means you could replace the vendor tomorrow. The firms that consistently deliver are the ones that treat strategy and build as one job, not two. If you would like to discuss a specific use case or roadmap for your organisation, AI Advisory works with mid-market UK businesses on exactly this pattern - discovery through to shipped systems, with a bias toward measurable output over decks.

Ready to put this into production? book a discovery call.

Get started

Ready to automate your operations?

Walk away with a prioritised list of automation and AI wins, costed, sequenced, and yours. The call is 30 minutes, free, and binds you to nothing. The shortest path to knowing whether AI Workflow Agency is the right fit.