Conversation
…ics and profiler Found by strict assertions in the BigQuery CLI E2E v2 migration (open-metadata/openmetadata-collate#5963): - bigquery: reflect NUMERIC/BIGNUMERIC as NUMERIC (was INT) and JSON as JSON (was VARCHAR); profile JSON with a pass-through type because the BigQuery dialect has no JSON deserializer. - bigquery system metrics: restore the per-table filter; every DML job in a dataset was attributed to every profiled table. - common db source: FK lookup falls back to the current database when a connector reports no referred_database (was `svc.None.schema.table`). - column handler: ARRAY/STRUCT/MAP columns get their NULL/NOT_NULL constraint like primitive columns. - bigquery sampler: pass SQLAlchemy 2 `_set_parent` arguments for STRUCT subfields. - profiler: unique count addresses STRUCT subfields by path and runs on the metric thread's session instead of the shared one.
…rk (#33947) Adds the `bigquery` connector suite (22 contracts) under ingestion/tests/cli_e2e_v2, covering every behaviour asserted by v1 test_cli_bigquery.py and test_cli_bigquery_multiple_project.py plus procedures, FKs, repeat ingestion and native samples. Each test owns a labelled, auto-expiring dataset in two GCP projects and scopes every workflow to it. Service-account auth is the CI default; E2E_BQ_AUTH=adc runs locally with Application Default Credentials. Offline meta-tests cover invocation scoping, credentials-by-reference and the pure checks. The v1 tests stay until the stability window completes. Refs open-metadata/openmetadata-collate#5963
❌ PR checklist incompleteThis PR cannot be merged until the following are addressed on its linked issue:
The fields live on the linked issue in the Shipping project (open the issue → right sidebar → Projects). After you set them, re-run this check (or push a commit) — issue/project changes do not re-trigger it automatically. Maintainers can bypass this check by adding the |
✅ Playwright Results — workflow succeededValidated commit ✅ 110 passed · ❌ 0 failed · 🟡 1 flaky · ⏭️ 0 skipped · 🧰 0 lifecycle flaky PerformanceBlocking targets: ✅ met · Optimization targets: 🟡 in progress Shard-job maxima below are not the full workflow wall time; the linked run includes build, fixture, planning, and reporting. 🕒 Full workflow signal wall (to summary) 50m 16s ⏱️ Max setup 4m 16s · max shard execution 12m 34s · max shard-job elapsed before upload 19m 45s · reporting 15s 🌐 106.68 requests/attempt · 1.81 app boots/UI scenario · 0.00% common-shard skew Optimization targets still in progress:
🟡 1 flaky test(s) (passed on retry)
How to debug locally# Download playwright-test-results-<shard> artifact and unzip
npx playwright show-trace path/to/trace.zip # view trace |
…STRUCT unique-count
Code Review ✅ Approved🟡 Medium risk · Shared ingestion fixes alter catalog types, constraints, and profiling across connectors Fixes six BigQuery ingestion and profiler bugs exposed by strict E2E assertions: NUMERIC/JSON columns now ingest as correct types, system metrics properly filter to per-table DML, foreign keys without OptionsDisplay: compact → Counting what did not apply, without listing it. Comment with these commands to change the behavior for this request:
Was this helpful? React with 👍 / 👎 | Powered by Gitar — free for open source |
|



Describe your changes:
Fixes open-metadata/openmetadata-collate#5963
I migrated the BigQuery CLI E2E coverage (
cli_e2e/test_cli_bigquery.pyandtest_cli_bigquery_multiple_project.py) to the v2 framework from #27949, because v2 asserts persisted OpenMetadata state strictly instead of v1's status-count floors. Those strict assertions, run against real BigQuery, exposed six connector/profiler bugs, which this PR also fixes (with a regression unit test for each). The v1 tests and theirpy-cli-e2e-tests.ymlmatrix entries stay until the agreed stability window completes. (Includes #33947, merged into this branch.)Type of change:
High-level design:
1. Connector and profiler fixes
Each fix is at the shared root cause rather than per caller:
INT, JSON →VARCHARbigquery/metadata.pyreflects them assqlalchemy.NUMERIC/sqlalchemy.JSON; newprofiler/orm/types/bigquery_json.py(BigQueryJSON, pass-through processors — the BigQuery dialect has no JSON deserializer and the DB-API already returns dicts) wired into the BigQuery ORM converter andNOT_COMPUTEsystem/bigquery/system.py: restore the per-table filter broken in #22044 (x or -1 > 0 and …, thennoqa: RUF021), same shape as Snowflakesvc.None.schema.tablecommon_db_source._prepare_foreign_constraintsfalls back to the current database whenreferred_databaseis absent, mirroring the existingreferred_schemafallbacksupportsDatabasethat doesn't report it (BigQuery, Databricks, Trino, Presto, Vertica, Synapse, DB2, Unity Catalog, …)sql_column_handler: complex branch calls_get_column_constraintslike the primitive branchColumn._set_parent() missing 'all_names'on STRUCT subfieldssampler/sqlalchemy/bigquery/sampler.py: SQLAlchemy 2 arguments, as #32549 did for the profiler interfaceKeyError: 'struct_col.x'; intermittent "session is provisioning a new connection"SQAProfilerInterface._compute_query_metrics: address a STRUCT subfield by path when it isn't a sample column, and build the grouped query on the metric thread'ssessioninstead of the sharedself.sessionRollout: on the next ingestion, existing BigQuery NUMERIC/JSON columns change data type, complex columns gain a constraint, and previously dropped FKs appear — one round of change events per affected column/table; nothing is removed.
2. BigQuery CLI E2E v2 suite
Follows
cli_e2e_v2/CONNECTORS.mdwith no shared-core changes. The new code is underingestion/tests/cli_e2e_v2/bigquery/:source.py,conftest.py): BigQuery can't run in a container, so each test creates its own datasete2e_bq_<uuid>(labelowner=cli-e2e-v2, 24h table expiration as a safety net) in two real projects, and deletes it with its contents on success and failure. Every invocation carries an anchoredschemaFilterPatternfor the owned datasets; a schema filter without explicit includes is rejected, so no run can ingest unowned project data.E2E_BQ_*env vars by reference (CI default).E2E_BQ_AUTH=adcruns locally with Application Default Credentials (gcp_adc).connector.py): single-project runs setbillingProjectIdto the second project, as v1 did; multi-project runs use aprojectIdlist.baseline.py,expected.py): authored independently of the connector's parsers, with an explicit BigQuery type map. Types are strict by review decision (NUMERIC →NUMERIC, JSON →JSON).tableDiff.Not included (needs maintainer authorization):
.github/workflows/py-cli-e2e-tests-v2.ymlmust allowlistbigqueryand pass the existingTEST_BQ_*secrets to the BigQuery job only;meta/test_ci_workflow.pyneeds the matching accepted case. Also note the job's 60-minute timeout: the suite takes about 9 minutes with-n 6and longer serially.Tests:
Use cases covered
NUMERICwith precision/scale, JSON asJSON; INT64/STRING unchanged.referred_databaseresolves in the current database.NULL/NOT_NULL.INFORMATION_SCHEMA.Unit tests
tests/unit/topology/database/test_bigquery_column_types.py,tests/unit/topology/database/test_complex_column_constraints.py,tests/unit/metadata/ingestion/source/database/bigquery/profiler/test_bigquery_system_profile.pytests/unit/topology/database/test_common_db_source.py,tests/unit/observability/profiler/sqlalchemy/bigquery/test_bigquery_profiler_sql.py,tests/unit/observability/profiler/sqlalchemy/bigquery/test_bigquery_sampling.pycoverage run -m pytestover the tests above):sampler/sqlalchemy/bigquery/sampler.py96%,orm/types/bigquery_json.py91%,orm/converter/bigquery/converter.py71%,metrics/system/bigquery/system.py68% (uncovered lines are the pre-existing JOBS query path; the changed filter is fully covered). Changed lines in the large shared files are covered by the new tests.tests/unit/topology/database,tests/unit/source/database, profiler/sampler suites — identical failures/errors to a cleanmaincheckout (all from optional connector packages missing locally), plus the new tests passing.ingestion/tests/cli_e2e_v2/meta/test_bigquery_cases.py(18 offline tests: owned-dataset scoping, credentials by reference, ADC config, invalid invocations, DQ config shape, strict type map, system-profile and partition checks).--e2e-contract-check --collect-onlycollects all 22 contracts exactly once.Backend integration tests
Ingestion integration tests
ingestion/tests/cli_e2e_v2/bigquery/: the live suite (real BigQuery → CLI → OpenMetadata) that found and verified the fixes.Playwright (UI) tests
Manual testing performed
open-metadata-beta+modified-leaf-330420with Application Default Credentials.E2E_BQ_AUTH=adc E2E_BQ_PROJECT_ID=open-metadata-beta E2E_BQ_PROJECT_ID2=modified-leaf-330420 python -m pytest ingestion/tests/cli_e2e_v2/bigquery --e2e-contract-check -n 6:catalog.multi-project) is a local IAM gap: my user lacks project-levelbigquery.tables.liston the second project, so the region-scoped lifecycle query gets a 403. The CI service account should not hit it.e2e_bq_*datasets ore2e_*services after the runs.UI screen recording / screenshots:
Not applicable.
Checklist:
I have read the CONTRIBUTING document.
My PR title is
Fixes <issue-number>: <short explanation>My PR is linked to a GitHub issue via
Fixes #<issue-number>above. — cross-repo link to the Collate tracking issue.I have commented on my code, particularly in hard-to-understand areas.
For JSON Schema changes: I updated the migration scripts or explained why it is not needed. — no schema changes.
For UI changes: I attached a screen recording and/or screenshots above. — no UI changes.
I have added tests (unit / integration / Playwright as applicable) and listed them above.
I have added tests around the new logic.
For connector/ingestion changes: I updated the documentation. —
ingestion/tests/cli_e2e_v2/README.md.