Insight · Microsoft Fabric

    Fabric Workspace Monitoring, Monitor Hub and Operations Agents: Observability in Practice

    Fabric workspace monitoring now centralises telemetry, the Monitor hub is rebuilt and an operations agent investigates failed pipelines. What to switch on, what it costs and what to watch.

    Nick de Vrye, CTOPublished 8 October 20266 min read
    Navy Solv Systems title card reading 'Fabric Workspace Monitoring' with a broadcast signal motif.

    In Short: Fabric Workspace Monitoring Becomes the Operations Layer

    Until now, finding out why something failed in Microsoft Fabric meant opening the failed item, then the next one upstream, then checking the capacity. At FabCon Europe 2026 Microsoft announced a joined-up observability layer to replace that routine. It has three parts, and all three are in preview:

    • Fabric workspace monitoring, rebuilt around a Monitoring Item that can send telemetry from many workspaces into one shared Eventhouse
    • A refreshed Monitor hub, with health views across jobs, apps, agents, alerts and capacities
    • Operations agent investigations, which read the telemetry for a failed pipeline run and suggest a likely root cause

    My view: switch on workspace monitoring for production workspaces now, centralise it from the start, and use the agent as a junior triage analyst rather than an oracle.

    How Fabric Workspace Monitoring Works Now

    You configure monitoring from Workspace settings > Monitoring, which creates a Monitoring Item. You then choose where the data goes:

    • Send data to Eventhouse in this Monitoring Item - provisions an Eventhouse with a read-only KQL database, an Eventstream and an Activator in the current workspace
    • Send data to Eventhouse in another Monitoring Item - sends this workspace's telemetry to an existing monitoring environment, which is how you centralise

    Data collection starts switched off, and Microsoft is clear that monitoring "doesn't backfill earlier activity". Turn it on before you need it. The supported telemetry already covers most of what a data platform team cares about: warehouse and SQL analytics endpoint query execution, pipeline activity runs, Copy job runs, Eventhouse query, command and ingestion logs (plus new throttling, sub-optimal size and scale-out logs), Eventstream metrics and errors, mirrored database execution, Activator rule notifications, GraphQL operations and semantic model operations.

    Everything lands in KQL tables, so your Eventhouse and KQL skills carry straight over. Microsoft publishes sample queries and Real-Time Dashboard and Power BI templates, and the Eventstream can be given a custom endpoint so external tools can consume the same events over Event Hubs, Kafka or AMQP.

    The constraints that shape the design

    • Only one Monitoring Item can collect telemetry for a workspace at a time
    • You cannot change the destination after configuration
    • Centralising requires both workspaces in the same Azure region
    • Each contributing workspace gets its own KQL database in the destination workspace, and those count towards the 1,000-item workspace limit
    • Data is retained for 30 days by default, adjustable through the database's retention and caching policies
    • Setup needs a workspace on Premium or Fabric capacity, the tenant setting allowing workspace admins to turn on monitoring, and the workspace admin role

    Because the destination is fixed, decide your topology before you click. That is a five-minute conversation now and a rebuild later if you skip it.

    The Topology We Recommend

    Microsoft's own guidance is to consolidate into as few Eventhouses as possible and to put the central monitoring workspace on its own capacity. We agree. The reasoning is in the documentation: isolation protects monitoring from throttling on busy production capacities, and protects production from the cost of the monitoring workload.

    For a typical mid-sized estate that means:

    • One monitoring workspace per region, on a small dedicated capacity
    • Every production workspace sending to it from day one
    • Development and feature-branch workspaces left out unless you are chasing a specific problem

    One detail matters for incident handling. Microsoft notes that monitoring ingestion, queries on the monitoring Eventhouse and Real-Time Dashboards over it keep working when the capacity is throttled, but Power BI reports and Activator alerts built on the monitoring database respect the capacity state and can be throttled. If your incident dashboard is a Power BI report on a shared, overloaded capacity, it may go quiet at the very moment you need it. Use a Real-Time Dashboard for the on-call view.

    If you are adopting deployment plans and branch workspaces, make monitoring setup part of the production workspace definition rather than a manual step someone remembers.

    The Refreshed Monitor Hub

    The Monitor hub was already the quickest place to check whether last night's jobs ran. The refreshed version (preview) adds health views across workloads, jobs, apps, agents, alerts and capacities. Capacity headroom now sits next to job status, and Capacity Insights and Actions (also preview) lets an administrator adjust surge and overage settings or move workspaces from the same screen. We covered those capacity controls in our capacity administration guide.

    The Job runs page lists Copy jobs, Dataflow Gen2, notebooks, pipelines, semantic models, Spark job definitions, user data functions and more. Dataflow Gen1 is not included. Alerting has two tiers, both in preview:

    • Failed scheduled job emails - free, and they cover scheduled runs only, not manual ones
    • Activator-based alert rules - billed at Activator rates, covering events such as Started and Succeeded as well as failures, and able to trigger automated actions. These need workspace monitoring enabled.

    Route the free failure emails to a monitored shared mailbox or group, not to whoever built the pipeline. The same principle applies to Power BI refresh failures: an alert that reaches someone on leave has not really been sent.

    Operations Agent Investigations

    This is the AI part. When a data pipeline run fails, you select Investigate in the Monitor hub and a chat opens with the operations agent. It reads the monitoring data for that run plus the pipeline definition, correlates the job, activity and error details, and returns a report: what happened, a timeline, the likely root cause and recommended next steps, with a run-history chart when there is enough history.

    The documented boundaries are reassuring:

    • It is read-only - it does not change the pipeline or rerun jobs
    • It runs as your identity, needs at least the Contributor role, and sees only what you can already see
    • The session is temporary and the conversation is not kept
    • It supports pipeline failures only for now
    • Microsoft states that billing isn't currently enabled for observability investigation

    Two practical catches. The investigation depends on workspace monitoring, so a run that failed before monitoring was switched on gives it nothing to work with. And if you enable the operations agent during monitoring setup, Microsoft says you cannot disable it later. Decide deliberately.

    A further preview, Investigator, extends the operations agent to analyse anomalies in business data as well as observability data. It is the same idea we described in real-time AI with Activator: detection by rules, context by AI, decision by people.

    What This Changes for Support Work

    For teams that run Fabric for others, internal or external, the effect is practical. With telemetry in one KQL database, a support runbook becomes a set of saved queries rather than a set of portal clicks. "Which pipelines failed overnight, which activity, which error class" becomes a single query across every workspace. That shortens time to diagnosis and it is the foundation of any sensible Power BI managed service.

    My checklist for a new production estate:

    • Agree the region-by-region monitoring topology before enabling anything
    • Enable workspace monitoring on every production workspace and confirm the tables populate
    • Set retention to match how far back incidents are actually investigated
    • Configure failure emails to a shared destination, and Activator rules only where an automated response is worth the consumption
    • Build the on-call view as a Real-Time Dashboard, not a Power BI report
    • Trial operations agent investigations on real failures and compare their conclusions with your own before relying on them

    Everything here is in preview, so check the current status on Learn before you build production processes around it. For the wider set of announcements, see our FabCon Europe 2026 recap.

    Where Solv Systems Comes In

    We design Fabric estates that can be run, not just built: monitoring topology, alert routing, runbooks and the capacity controls that go with them. If your current alerting is a person checking the Monitor hub each morning, we can help you replace it.

    Sources and Further Reading

    Frequently asked

    It is the optional observability layer for a Fabric workspace, currently in preview. You configure a Monitoring Item, which provisions an Eventhouse with a read-only KQL database, an Eventstream and an Activator. Supported items such as pipelines, Copy jobs, warehouses, Eventhouses, Eventstreams, mirrored databases and semantic models then send logs and metrics you can query with KQL.

    Yes. When you configure monitoring you can send data to the Eventhouse in another Monitoring Item, which centralises telemetry. Both workspaces must be in the same Azure region, the choice cannot be changed later, and each contributing workspace gets its own KQL database in the destination workspace, which counts towards that workspace's 1,000-item limit.

    Yes. Microsoft states that workspace monitoring uses Fabric capacity, billed through the Eventhouse, KQL database and Eventstream consumption rates. Microsoft recommends hosting the central monitoring Eventhouse on its own capacity so monitoring and production workloads do not compete.

    It is a read-only, AI-assisted investigation of a failed data pipeline run, started from the Monitor hub. It correlates job, activity and error details, produces a report with a likely root cause and next steps, and accepts follow-up questions. It is in preview, supports only pipeline failures today, and Microsoft says billing is not currently enabled for it.

    30 days by default. You can change the retention and caching policies on the monitoring KQL database, with caching set to no longer than retention. Monitoring does not backfill, so anything that happened before collection was switched on is not captured.

    The email notifications for failed scheduled jobs are free. Activator-based alert rules, which cover more job events such as Started and Succeeded and can trigger automated actions, are billed at Activator consumption rates and need workspace monitoring enabled. Both are in preview.