Data lineage graph
Latest stable Based on Graph release 0.5.1
Problem and backend
Section titled “Problem and backend”This example follows upstream sources and downstream impact across datasets, jobs, and reports. It uses TinkerGraph to isolate modeling from container and network variance. Read core model and TinkerPop first; use the selection guide before production.
- Nodes: Dataset/Table/Column/PipelineJob/Dashboard/Owner/QualityCheck
- Edges: CONTAINS_TABLE/CONTAINS_COLUMN/INPUT_TO_JOB/OUTPUTS_TABLE/FEEDS_DASHBOARD/OWNS_JOB/VALIDATES_COLUMN
- Key properties: datasetId, tableId, columnId, jobId, dashboardId, ownerId, checkId
Prerequisites and release boundary
Section titled “Prerequisites and release boundary”Use JDK 21, commit 3e0fa7cb9e3bc70c2743aeebda2487f3e45e4907, and the checked-in wrapper. Examples are not published; run this release fixture as a Gradle project from the release source checkout. In a consumer application, select only bluetape4k-dependencies:<ecosystem-version> and add the required graph module without an individual version.
Run and observe
Section titled “Run and observe”./gradlew :data-lineage-examples:test --tests "io.bluetape4k.graph.examples.datalineage.TinkerGraphDataLineageImpactTest"The test asserts that downstream impact reaches exec-revenue and ops-quality and that upstream traversal finds the expected source tables. A failure points first to lineage direction, missing transformation edges, or a changed traversal bound.
Reading order
Section titled “Reading order”Continue from observability-graph, then read supply-chain-graph. Also see paired APIs, testing, and operations.
Exercises and production gaps
Section titled “Exercises and production gaps”Add one result-changing edge and assertion; repeat through the suspend API; then run a persistent-backend concrete test serially. Add disconnected and malformed inputs as diagnostics. This fixture does not prove throughput, clustering, authorization, tenant isolation, migration, backup, remote-driver timeout, or index quality.