Repository navigation
sqlmesh format should not fully load the project when paths are given #6025
Description
Activity
I'd like to take this one if it's free — unassigned with no linked PR right now.
Plan: skip
Context.load()when positional paths are given and format those files directly._formatonly needs three things from the loaded object —_path,dialectandformatting— and all three are available from the file's ownMODEL/AUDITheader plusconfig_for_path, which resolves a project config from a path without loading anything.is_meta_expressionon the first parsed expression is what separates a model or audit file frommacros/*.sql, so non-model SQL stays a no-op without needing the project graph.One behavioural question before I start. Today "is this file a model?" is answered by membership in the loaded project. Without the load it would be answered by parsing the header, and those two differ for a
.sqlfile that has a MODEL header but is excluded from the project — sitting outsidemodels/, or matched byignore_patterns. Todaysqlmesh format path/to/that.sqlskips it; a pure header check would format it.I'd rather preserve today's behaviour and still honour
ignore_patternsand themodels//audits/directories, since both are cheap to check against the config without loading. Say if you'd prefer the simpler "any file with a MODEL/AUDIT header gets formatted" instead.Also flagging that #5944 touches
Context.format()as well. It's been quiet since August — I'll keep this change confined to the load path so the two stay separable.One finding while mapping this out, since it affects the spec above.
The acceptance says to format a path "if [it is] a SQL model or standalone audit", but standalone audits aren't formatted today.
Context.formatiterateschain(self._models.values(), self._audits.values()), and anAUDIT (..., standalone true)is loaded intoself._standalone_audits, which is a separate dict. So it's skipped.Reproduced on
mainwith a project holding one model and two audit files:audits/ma.sql (AUDIT(name ma, dialect 'duckdb')) -> reformatted audits/sa.sql (AUDIT(name sa, dialect 'duckdb', standalone true)) -> byte-identical, untouchedSo "format only those paths that are models or standalone audits" is a behaviour change on top of the load change, not just a restatement of what happens now.
Happy to include it — the header-based check I described treats both kinds of
AUDITthe same, so standalone audits would start being formatted, which I think is what you want given the wording. But it's a separate change in effect, so tell me if you'd rather I preserve the current skip and let the standalone-audit gap be its own issue. I'll keep it isolated in the diff either way.
Summary
sqlmesh formatalready accepts positional paths and does not load state. With paths it still fully loads the project, then formats only the matching.sqlmodels and audits. That load is wasted: format pretty-prints file text using dialect and format config. It does not need other models.Current behavior
Constructs
Contextwithload=True, loads every model, then filters withPath.samefile.Proposed behavior
Context.load()when positional paths are presentformatting falseformat:from the file’s project config and MODEL/AUDIT DDLmacros/*.sqland other non-model SQL: no-op (same as today)Acceptance
sqlmesh format models/a.sqldoes not parse unrelated modelssqlmesh formatwith no args is unchanged