Skip to content

fix: quote struct field names that need it in column types - #1710

Open
maharanay22 wants to merge 1 commit into
databricks:mainfrom
maharanay22:fix-struct-field-name-quoting
Open

maharanay22 wants to merge 1 commit into
databricks:mainfrom
maharanay22:fix-struct-field-name-quoting

Conversation

@maharanay22

Copy link
Copy Markdown

No existing issue (I searched open and closed issues and PRs; #1070 is a different struct problem caused by truncated types).

Description

DatabricksColumn._parse_type_from_json builds struct types from DESCRIBE ... AS JSON without quoting field names, so a field like first name or order-id produces struct<first name:string,order-id:bigint>. That is invalid DDL, so the materialization V2 CREATE and on_schema_change ADD COLUMNS fail with PARSE_SYNTAX_ERROR for tables with JSON-style field names.

Field names are now quoted only when Spark's QuotingUtils.quoteIfNeeded would quote them: names matching [A-Za-z_][A-Za-z0-9_]* stay bare, and anything else is wrapped in backticks with embedded backticks doubled. Type strings for ordinary field names are unchanged (see #1148 for why that matters). Nested structs inside arrays and maps are handled by the existing recursion.

New parametrized tests cover plain names (unchanged), space, hyphen, dot, colon, an embedded backtick, a leading digit, and a struct nested in an array and a map. 7 of the 8 fail on main (the plain-name case passes by design) and all pass with this change.

Checklist

  • I have run this code in development and it appears to resolve the stated issue
  • This PR includes tests, or tests are not required/relevant for this PR
  • I have updated the CHANGELOG.md and added information about my change to the "dbt-databricks next" section.
  • [Optional] I have run the dbt-databricks-pr-ready project skill for this PR and addressed its merge-readiness feedback

`DatabricksColumn._parse_type_from_json` wrote struct field names from
DESCRIBE ... AS JSON without quoting, so a field like `first name` or
`order-id` produced `struct<first name:string,order-id:bigint>` and the
materialization V2 CREATE or the on_schema_change ADD COLUMNS failed
with PARSE_SYNTAX_ERROR.

Quote a field name only when Spark's quoteIfNeeded would, backticking
it and doubling embedded backticks, so type strings for ordinary names
do not change. Nested structs inside arrays and maps are covered by the
existing recursion.

Signed-off-by: Maha Rana Yadavalli <271375718+maharanay22@users.noreply.github.com>
@maharanay22
maharanay22 force-pushed the fix-struct-field-name-quoting branch from f11093c to 8bb4f5a Compare October 8, 2026 22:35

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant