Fabric Semantic Model Refresh: One Table at a Time

In a previous post I showed how to trigger a semantic model refresh from a Fabric notebook using sempy.fabric. That approach works great for small models, but in production you’ll quickly run into problems when the model has many tables or large fact tables. The problem: out of memory The sempy default for max_parallelism is 10. For a model with 27 tables including several large fact tables (tens of millions of rows), processing 10 tables concurrently can easily exhaust the available memory on the capacity and the refresh fails with a generic “out of memory” error — or worse, it succeeds but takes an unreasonably long time because of internal retries. ...

October 7, 2026 · 6 min · Pedro Morais

Microsoft Fabric External Data Sharing Hidden Limitation

Microsoft Fabric external data sharing is a powerful feature that allows us to share data with external users. This is particularly interesting because it enables data to be shared in-place from the sharer’s tenant without any data copy. We simply create shortcuts. The use case for the customer in question involves data from Microsoft Dynamics 365 (D365) residing on two different Azure tenants. I needed a solution to share data from one tenant to the other. ...

April 9, 2026 · 3 min · Pedro Morais

Refreshing Power BI Datasets with Notebooks

The sempy Python library in Fabric is quite powerful and offers a range of capabilities that make working with Fabric environments more efficient and flexible. One of the most useful features I’ve implemented in my projects is the ability to trigger semantic model refreshes directly from a Spark notebook. This approach allows to trigger the refresh right after the data load is finished. import sempy.fabric as fabric import time workspace = fabric.resolve_workspace_name() dataset = "My_Amazing_Semantic_Model" refresh_request_id = fabric.refresh_dataset(workspace=workspace, dataset=dataset, refresh_type="full") # refresh api is async, so we need to poll until it completes while True: time.sleep(60) res = fabric.get_refresh_execution_details( dataset=dataset, workspace=workspace, refresh_request_id=refresh_request_id ) if res.status != "Unknown": break if res.status == "Completed": print(f"Refresh Dataset completed with success {res.extended_status}") else: errors = res.messages["Message"].str.cat(sep="\n") raise Exception(f"Refresh Dataset failed: \n\n{errors}") This script starts by resolving the current workspace and identifying the target dataset. It then initiates a full refresh of that dataset and waits for the process to complete, checking periodically for updates. ...

November 12, 2025 · 2 min · Pedro Morais

Terraform/OpenTofu Notes

Using Workspaces for Environment Separation Terraform workspaces are a powerful feature that allows you to manage multiple environments (such as dev, staging, and prod) within the same configuration. For example, to switch to the production workspace: tofu workspace select prod Applying Configuration Changes When applying changes to your Terraform configuration, you can specify a variable file for different environments and approve the changes automatically: tofu apply -var-file ./dev.tfvars -auto-approve Importing Existing Resources To import an existing resource into Terraform’s state, use the import command with the appropriate variable file: ...

November 5, 2025 · 3 min · Pedro Morais

The credentials provided for the UsageMetricsDataConnector source are invalid.

Last weeks I had several meetings with Microsoft support to try this wrong credentials issue. The problem was that after right clicking to “View Usage Metrics Report” the report would come up empty and after 24h you would get an email with the following message: <ccon>The credentials provided for the UsageMetricsDataConnector source are invalid. (Source at UsageMetricsDataConnector.)</ccon>. The exception was raised by the IDbCommand interface We tried several approaches until we found this piece of documentation: https://learn.microsoft.com/en-us/power-bi/collaborate-share/service-modern-usage-metrics#considerations-and-limitations ...

July 14, 2025 · 1 min · Pedro Morais

Issues with v4l2loopback and OBS

I started my Linux journey a long time ago with Red Hat in the middle of the nineties and moved after to Slackware. Still to this day one of my favourite distros. I tried Ubuntu for a shortwhile but found a new home on Arch. It was my main distro for the better part of the 2010’s but then tried Void Linux and it’s now my home since at least 3 years. ...

July 9, 2025 · 4 min · Pedro Morais

Tabular Editor Scripts: Simplifying Your Power BI Modeling Experience

As a data enthusiast, I’m always on the lookout for ways to streamline my workflow and make my life easier. In this post, I’ll be sharing two Tabular Editor scripts that have saved me countless hours of manual formatting in my Power BI projects. There are two versions of this software. To use the scripts below you only need version two. This version is free and open source: https://tabulareditor.github.io/TabularEditor/ If you have the money for the license you can support the author (Daniel Otykier) by buying the license for version 3: https://tabulareditor.com ...

April 27, 2025 · 3 min · Pedro Morais

Extracting tables from SQL queries by using Sqlfluff

I got an assignment the other day to produce documentation to send to a customer. The extraction of the table names required to execute a certain Databricks notebook was part of the task. The plan was to build an object dependency tree. The query spanned 279 lines. How can you extract only the table names from a file without having to manually look for them? Can we make use of this technique again in the future? ...

September 3, 2023 · 3 min · Pedro Morais

Azure Active Directory extraction with Databricks

During data engineering projects I tend to try and minimize the tools being used. I think it’s a good practice. Having too many tools causes sometimes for errors going unnoticed by the teams members. One of the advantages of having a tool like Databricks is that it allows us to use all the power of python and avoid, like I did in the past, to have something like Azure Functions to compensate for the limitations of some platform. ...

May 8, 2023 · 2 min · Pedro Morais

List of errors from Databricks API

I’m currently working on a project where I’m adapting a code base of Databricks notebooks for a new client. There are a few errors to hunt but the Web UI is not really friendly for this purpose. Just wanted a quick and easy way to not have to click around to find the issues. Here’s a quick script to just do that: import os, json import configparser from databricks_cli.sdk.api_client import ApiClient from databricks_cli.runs.api import RunsApi def print_error(nb_path, nb_params, nb_run_url, nb_error="Unknown"): error = nb_error.partition("\n")[0] params = json.loads(nb_params) if nb_params != "" else {} print( f""" Path: {nb_path} Params: {json.dumps(params,indent=2)} RunUrl: {nb_run_url} Error: {error} """ ) databricks_cfg = "~/.databrickscfg" conf = configparser.ConfigParser() conf.read(os.path.expanduser(databricks_cfg)) api_client = ApiClient( host=conf["DEFAULT"]["host"], token=conf["DEFAULT"]["password"] ) runs_api = RunsApi(api_client) for x in range(1, 101, 25): x = runs_api.list_runs( job_id=None, active_only=None, completed_only=None, offset=x, limit=25, version="2.1", ) if len(x["runs"]) > 0: for y in x["runs"]: if y["state"]["result_state"] == "FAILED": z = runs_api.get_run_output(run_id=y["run_id"]) if "error" in z: print_error( z["metadata"]["task"]["notebook_task"]["notebook_path"], z["metadata"]["task"]["notebook_task"]["base_parameters"][ "Param1Value" ], z["metadata"]["run_page_url"], z["error"], ) else: print_error( z["metadata"]["task"]["notebook_task"]["notebook_path"], z["metadata"]["task"]["notebook_task"]["base_parameters"][ "Param1Value" ], z["metadata"]["run_page_url"], ) Follow this documentation to install the requirements. There’s a lot more you can do with databricks-cli to make your life easier. It’s a great tool to add to your toolbox. ...

February 17, 2023 · 1 min · Pedro Morais