Build and operate the analytics solution
- Set up Fabric workspaces: Spark, domains, OneLake and Apache Airflow
- Manage the lifecycle: version control, database projects, deployment pipelines
- Security and governance: access at workspace and item level, row, column and object level, dynamic data masking, sensitivity labels, audit logs, OneLake security
- Orchestrate processes: choosing between Dataflow Gen2, pipeline or notebook; schedules and event-based triggers; orchestration patterns with parameters and dynamic expressions
Ingest and transform data
- Design load patterns: full and incremental loads, preparation for a dimensional model, load patterns for streaming data
- Batch data: choose the right data store, transformation with Dataflows Gen2, notebooks, KQL or T-SQL
- OneLake shortcuts and mirroring, ingestion via pipelines
- Transformation with PySpark, SQL and KQL: denormalise, group, aggregate, and handle duplicates as well as missing and late-arriving data
- Streaming data: choose the right streaming engine, eventstreams, Spark Structured Streaming, KQL, window functions
Monitor and optimise
- Monitor ingestion, transformation and semantic model refresh, set up alerts
- Narrow down and resolve errors – in pipelines, dataflows, notebooks, Eventhouse, eventstream, T-SQL and OneLake shortcuts
- Optimise performance: lakehouse tables, pipelines, data warehouse, eventstreams and eventhouses, Spark and queries