Insight · Microsoft Fabric

    What Copilot in Fabric Actually Costs: Capacity Consumption Explained

    Copilot in Fabric has no per-user licence: every prompt bills capacity units against your F SKU. How the consumption model works, and how to keep it from eating your capacity.

    Nick de Vrye, CTOPublished 7 September 20266 min read read
    Navy Solv Systems title card reading 'What Copilot in Fabric Costs' with a capacity gauge motif.

    In Short: The Meter Is the Capacity

    Copilot in Fabric and Power BI is charged the way Fabric charges everything: consumption of capacity units against your F SKU. There is no per-user licence, no separate AI bill, and no obvious price tag on any single prompt, because the cost is folded into the same capacity that runs your refreshes, queries and pipelines.

    That model is convenient and dangerous in equal measure. Convenient because enabling Copilot involves no procurement. Dangerous because AI consumption arrives silently inside an existing budget line, and on smaller capacities it competes directly with the workloads your business already depends on.

    From F64-Only to Every Paid SKU

    When Copilot in Fabric first shipped, it required an F64 capacity or larger, a threshold that excluded most mid-market estates for whom F64 represents a substantial annual commitment. That gate has since been opened: Copilot can now be enabled on paid F SKUs from F2 upwards, with consumption metered against their capacity units.

    The relaxation changed the question from "can we afford Copilot" to "can our capacity absorb Copilot", which is a better question but not a trivial one. A capacity sized precisely for its refresh and query load, per our capacity sizing guide, has no headroom earmarked for conversational AI, and Copilot consumption is burstier and less predictable than scheduled workloads.

    How the Metering Works

    Mechanically, Copilot operations consume tokens - the prompt going in, the response coming out, and the model context in between. Fabric converts that token usage into capacity unit seconds and accounts for it as a background operation, smoothed over time like other background work rather than spiking the moment a prompt runs.

    Three practical consequences follow.

    • Visibility exists, but you must look. The Fabric Capacity Metrics app itemises Copilot consumption by operation, workspace and user. Nobody reviews it by default
    • Longer context costs more. Copilot over a large, complex semantic model processes more context per question than over a focused one; model hygiene has a direct cost dimension
    • Cost per prompt is invisible to users. The person asking questions sees no meter, so usage grows on convenience alone. Governance has to come from capacity monitoring, not user restraint

    Keeping Copilot From Eating the Capacity

    The estates that run Copilot comfortably do four things.

    • Enable it deliberately: pilot workspaces first, not tenant-wide on day one
    • Review the Capacity Metrics app weekly during rollout and monthly thereafter, watching Copilot's share of total consumption
    • Set expectations by role: viewer Q&A on governed models is cheap per interaction; heavy authoring assistance and long conversational sessions are not
    • Consider a dedicated Copilot capacity once AI usage is material, so AI bursts are isolated from production refreshes and the finance team gets a clean AI cost line

    If your capacity already runs warm, fix that before enabling Copilot; our work on Fabric cost measurement covers the diagnostic side. Adding conversational load to a capacity at the edge of throttling converts an AI rollout into a reliability incident.

    What This Means for Budgeting

    Because consumption depends on prompt volume, model size and conversation habits, published projections are close to worthless. The credible sequence is: enable for a pilot group, measure four to six weeks of real consumption in the Capacity Metrics app, extrapolate to the rollout population, and only then decide whether the existing SKU absorbs it or the budget needs a capacity step.

    Treat the resulting number the way you treat refresh load: an operating cost of the data platform, reviewed on the same cadence, owned by the same person. Copilot in Fabric is not a licence decision; it is a capacity planning decision, and the organisations that treat it that way avoid both the overspend and the throttling stories. The wider licensing picture across the Copilot family is mapped in our Copilot licensing guide.

    Sources and Further Reading

    Frequently asked

    No. There is no per-user Copilot licence for Fabric or Power BI. Copilot usage is metered as capacity unit consumption against a paid Fabric capacity that an administrator has enabled for Copilot. The cost appears in your capacity, not your licence bill.

    No. The original release required F64 or larger, which put Copilot out of reach for most mid-market estates. Microsoft subsequently opened Copilot to all paid F SKUs, so smaller capacities can enable it, with consumption billed against their capacity units. Confirm the current floor in Microsoft's documentation.

    Copilot operations convert the tokens your prompts and responses consume into capacity unit seconds, which are smoothed and reported like other background operations. The Fabric Capacity Metrics app breaks out Copilot consumption so you can see exactly which workspaces and users generate it.

    Yes, in the same way any workload can: capacity consumption is shared, so heavy Copilot use on a small SKU competes with refreshes, queries and pipelines. If the capacity sustains overload, everything on it degrades. Monitoring Copilot's share early tells you whether to budget more capacity.

    An option for attributing AI consumption: organisations can designate capacity specifically for Copilot billing, separating AI costs from core BI workloads. It is worth considering once Copilot usage is material, both for chargeback clarity and to stop AI bursts affecting production refreshes.

    Enable it for a pilot group, run four to six weeks, and read the Capacity Metrics app. Real prompt volumes vary so widely between organisations that projections are guesses; measured consumption from your own users is the only budgeting input worth having.