Harness Platform
Blog →
Harness Platform

Knowledge Graphs for Software Delivery: An Architectural Approach | Harness Blog

Discover how software delivery knowledge graphs unify fragmented SDLC data, enable schema-as-code, and power deterministic AI reasoning across engineering teams.

TL;DR

Software delivery knowledge graphs unify fragmented SDLC data across disparate tools by establishing explicit entities, typed relationships, and schema-as-code contracts. This transforms manual cross-system data stitching and non-deterministic AI reasoning into reliable, queryable knowledge.
Key takeaways include:

  • Core Architecture: Combines schema-as-code, unified synchronization, and declarative queries to make meanings and connections explicit.
  • Five Pillars of Trust: Relies on shared semantics, explicit contracts, high data quality (fill rates), detailed provenance for lineage tracking, and automated governance.
  • Operational Discipline: Treating data infrastructure with application-level rigor, using active monitoring, continuous evaluation, and production trace observability to prevent silent data degradation.

Engineering teams are drowning in data but starving for knowledge—the same entity has different identifiers across systems, cross-domain queries require manual stitching, and AI agents can't reason deterministically. 

Knowledge graphs solve this by making entities, relationships, and semantics explicit and queryable through schema-as-code, unified sync, and declarative queries. The critical insight: graph quality equals data quality. Without operational discipline, monitoring fill rates, tracking provenance, continuous evaluation—even well-designed graphs degrade into unreliable answers.

The Question That Started Everything

"Why did this deployment fail?"

It's 2 AM. An engineer is debugging a production incident. The deployment failed, but the error message is cryptic. She needs to understand: what changed in the code, which build packaged it, what security scans ran, which artifact was deployed, and what infrastructure it touched.

Every piece of data exists. In six different systems. With six different identifiers. No shared schema. No explicit connections.

She spends 15 minutes querying CI logs, artifact registries, security scanners, deployment tools, and observability platforms. Each system has its own API, its own naming conventions, its own view of reality. She stitches the story together manually, in a spreadsheet.

Ask the same question an hour later, query systems in a different order, and get different pieces of data. The answer changes.

Engineering teams are drowning in data but starving for knowledge.

The Knowledge Gap in Software Delivery

The problem isn't volume. It's connections.

Your CI system calls it payment-service. Your deployment system calls it payments. Your observability platform calls it payment-api. Your cost system calls it payment-svc. Four identifiers, one logical service.

When an AI agent tries to answer "Why did this deployment fail?" it faces the same problem the engineer faced: orchestrate multiple API calls, guess which fields join, handle missing data, and stitch responses. Non-deterministic. Ask twice, get different answers.

This isn't a data collection problem. It's a knowledge problem.

The data exists. The connections don't.

What is a Knowledge Graph?

Before diving into architecture, let's establish what we're building.

A knowledge graph is a structured representation of entities and their relationships. Unlike databases that store rows and columns, or data lakes that store files, knowledge graphs store meaning—what entities are, how they connect, and what those connections represent.

Think of it as the difference between:

  • A phone book (data): Names and numbers in alphabetical order
  • A social network (knowledge): People, their relationships (friend, colleague, family), shared interests, mutual connections

Context engineering is the discipline of building, maintaining, and operating these graphs for AI systems. It's not just data engineering (move data from A to B) or MLOps (train and deploy models). It's the systematic work of connecting data so that reasoning systems—human or AI—can query across domains.

Three properties distinguish knowledge graphs from other data systems:

1. Entities are first-class. Not just rows in a table. Each entity has an identity, a type, and meaning. A "service" isn't {id: 123, name: "payments"}—it's a concept with lifecycle, ownership, and dependencies.

2. Relationships are explicit. Not inferred from naming conventions or foreign keys. A deployment → artifact relationship is declared, typed, and enforced. The graph knows how things connect.

3. Semantics are shared. "Artifact" means the same thing to your build system, deployment system, and security scanner. Shared definitions enable cross-domain queries.

Why this matters for software delivery: your engineering data spans dozens of systems. Without a knowledge graph, each system is an island. Questions that span systems require humans to manually stitch data. With a knowledge graph, the connections exist and are queryable.

entity-chain-diagram.png

Architecture: How Knowledge Graphs Work

Here's the end-to-end topology:

architecture-diagram.png

 Data flows top to bottom. Each layer transforms raw data into structured knowledge.

What We Built: From Data to Knowledge

A knowledge graph solves this by making three things explicit: what entities mean across boundaries, how they connect, and whether you can trust the connections.

When these hold, queries become deterministic. Same question, same answer, every time.

But building it required rethinking how engineering teams model, share, and govern data. Here's the architecture we landed on.

Layer 1: Schema as Code

First principle: entity schemas are public contracts.

Just as breaking an API endpoint impacts downstream consumers, breaking an entity schema impacts downstream queries. When deployment entities reference artifacts through artifact_digest, every query joining deployments to artifacts depends on that field existing and meaning the same thing tomorrow.

So schemas live in version control. Not documentation. An executable schema that query compilers read at runtime.

Try to query a non-existent field? Error before execution. Try to join the wrong field? Error. The type system enforces correctness.

Ownership is decentralized. Domain teams own their schemas. But cross-domain relationships require approval from both sides. One team can't unilaterally create dependencies on another's entities.

This isn't bureaucracy. It's the same discipline you already apply to API versioning: deprecation notices, backward compatibility windows, coordinated rollouts.

Layer 2: Unified Synchronization

Entities flow from source systems into a unified query layer.

The synchronization layer does the following things: validates against schema on write (reject malformed data), enforces referential integrity (verify references resolve), derives relationships automatically (when valid references appear, create edges).

Choose a synchronization strategy by use case. "Why did this fail right now?" demands sub-second freshness. "What was our success rate last quarter?" tolerates the batch.

The key: keep source data immutable. Don't merge vendor entities destructively. The unified layer is a materialized view, recomputable with better algorithms.

Layer 3: Declarative Queries

Queries declare what to retrieve, not how to join.

Validation before execution. Registered relationships eliminate guessing. Authorization filters before returning. Whether SQL-like, GraphQL, or domain-specific, the principle holds: queries against schema, not implementation.

This is what makes AI agents viable. LLMs can generate valid queries because the schema is explicit and queryable. No tribal knowledge required.

Five Pillars of Trust

Architecture without operations is fiction. We learned that making the graph trustworthy required five operational disciplines.

1. Shared Semantics

"Artifact" must mean the same thing to your build system, deployment system, security scanner, and supply chain tool. Same definition, same identifier, same lifecycle.

Without this, cross-domain queries return nonsense. A security query that joins "artifacts" from two systems that define artifacts differently produces garbage.

This isn't a documentation problem. Documentation drifts. This is a contract problem. Enforce it at write time.

2. Explicit Contracts

Data contracts declare how entities connect: artifacts identified by content digest, produced by one build, deployed by zero or more deployments, scanned by zero or more security checks.

Contracts specify relationship keys, cardinality, and lifecycle rules. When the security team wants to relate scans to artifacts, the artifact team reviews. Both sides own the relationship.

Treat schemas with API-level rigor. A field rename breaks every query using it. Schema changes need the governance you already apply to APIs.

3. Data Quality = Agent Quality

Here's the insight that changed how we thought about this: schema completeness doesn't equal data completeness.

You can define artifact.digest in the schema. But if only 60% of artifacts have digests populated, 40% of the knowledge graph is invisible to queries. An agent asked, "Show me security risks in production deployments," can only be as accurate as the deployment → artifact → scan relationship fill rate.

Track four dimensions: completeness (percentage populated), correctness (matches reality), freshness (write-to-availability latency), consistency (references resolve).

Alert on degradation. Each domain team owns quality for their entities. Data quality is an operational concern, not a nice-to-have.

4. Provenance: Trust Through Lineage

Every entity in the graph carries its history: where it came from, when it arrived, how it was derived.

Provenance metadata includes:

  • Source system: Which operational database wrote this entity
  • Source timestamp: When the write happened in the source
  • Ingestion timestamp: When it entered the graph
  • Derivation lineage: If this entity was derived from others, which entities and which rules

Why provenance matters:

Debugging data quality. A deployment entity is 10 minutes stale. Provenance shows: CDC pipeline from MongoDB → Kafka topic lag is 8 minutes → processing took 2 minutes. Now you know where the delay is.

Auditing relationships. A deployment → artifact relationship exists. Who created it? Was it automatically derived (high confidence) or manually linked (requires verification)?  Provenance shows: created by reconciliation service with confidence score 0.92 based on container image match.

Re-deriving with better logic. You improve the reconciliation algorithm. Can you recompute all relationships? Yes—source data is immutable, provenance shows which entities were derived. Rerun derivation, update provenance, republish.

Compliance and governance. Which team owns this data? When was it last updated? What's the authoritative source? Provenance answers these without querying original systems.

Provenance isn't optional at scale. When you have hundreds of entity types flowing from dozens of systems, the question "where did this data come from?" becomes critical. Immutable source data + versioned provenance = debuggable, trustworthy knowledge.

5. Automated Governance

Schema evolution is inevitable. Fields get renamed. Relationships change. Types split or merge.

Without governance, these changes break downstream consumers silently. One team renames a field. Another team's queries still reference the old name. The query succeeds syntactically but returns empty results. Nobody notices until users complain.

Governance isn't approval boards. It's automated protection:

  • Contract tests that must pass
  • Evaluation harnesses that maintain accuracy
  • Monitors that stay green

Teams move quickly. Governance runs in CI, not in meetings.

The Obstacles We Hit

Four production challenges emerged that we hadn't anticipated.

Entity Reconciliation: The Identity Crisis

The same service appears as payment-api in Datadog, payments in your CD system, PaymentAPI in observability, and payment-svc in cost data.

How do you know they're the same entity?

We use composite keys for stable identity, declare golden sources for conflicts, and apply fuzzy matching with confidence scores. High-confidence matches auto-accept. Low confidence requires human review.

The stakes: incorrect matches break relationships. A query like "which team owns this cost spike?" returns the wrong answer if service reconciliation is wrong. Cross-domain queries fail silently—they return results, just incorrect ones.

Schema Evolution: The Coordination Problem

You have 30+ domains evolving independently. How do you coordinate schema changes without creating bottlenecks?

The pattern: decentralized ownership, centralized contracts. Domain teams evolve schemas freely. Cross-domain relationships require both sides' approval. Breaking changes get deprecation windows.

Graph Deterioration: The Silent Failure

Data quality degrades without announcing itself.

A synchronization pipeline breaks. No alerts fire (the service is running, just not processing events). New entities stop flowing. Queries return increasingly stale results. Users notice something is wrong eventually, but can't tell what's stale.

Unlike application bugs that fail loudly, data problems are invisible until someone queries the affected path.

We learned: active monitoring is non-negotiable. Track entity counts (sudden drops indicate pipeline failures), relationship coverage (percentage with required links), query health (latency, timeouts, and empty results).

Run automated hygiene continuously: prune orphaned edges, archive old entities, recompute derived relationships.

Treat data infrastructure with application-level rigor: SLOs, monitors, runbooks, and incident response.

Continuous Evaluation: The Trust Problem

How do you know the graph helps AI agents? Not "queries are fast." Does it improve task success?

We built an evaluation harness: a golden dataset of cross-domain questions with ground-truth answers. Compare graph-based retrieval against alternatives. Measure correctness, determinism, and latency.

The surprising finding: determinism matters as much as correctness.

A system that produces the right answers 80% of the time but gives different results on repeated queries? Users stop trusting it. They can't tell when it's right. They assume it's always wrong.

We run evaluations on every schema change. Accuracy regression below the threshold blocks the change.

But evaluation harnesses are synthetic: They test the questions you thought to ask. Real users ask questions you didn't anticipate.

So we instrument every query with observability traces. Every question asked, every HQL query generated, every relationship traversed, every entity accessed, captured, and analyzed.

This reveals:

What questions users actually ask: Not what we thought they'd ask. The difference teaches us which relationships to prioritize, which entity types need better coverage.

Which queries succeed vs. fail: Failed queries expose gaps. A query joining deployments to artifacts fails for 40% of deployments? The artifact relationship has a 60% fill rate. Now we know where data quality is failing.

Which relationships are missing: Users ask questions that span domains we haven't connected yet. "Which services depend on this database?" requires a service → database relationship we haven't modeled. Usage patterns drive schema evolution.

Which entities are incomplete: Queries filtering by owner_team fail when that field is unpopulated. Trace analysis shows exactly which entity types have coverage gaps.

This continuous feedback loop—production traces informing schema evolution—is how the graph gets better over time. Not from roadmap planning. From observing what breaks.

Use Cases: From Impossible to Automatic

The graph doesn't just speed up queries; it makes certain questions answerable at all, because the connections that answer them now exist.

DORA Metrics. Lead time, deployment frequency, change failure rate, and MTTR require joining Git, CI, artifacts, and deployments, normally stitched together by custom scripts on a schedule. A single query following commit → build → artifact → deployment makes this fast enough to run continuously. Enables real-time DORA tracking instead of periodic reports.

Security Response. When a vulnerability is disclosed, teams need to know which production services are affected, normally a manual cross-reference across SBOMs, artifacts, and deployments. A single traversal (component → SBOM → artifact → deployment → service) returns every match with ownership context. Enables precise, fast notification instead of broad "you might be affected" alerts.

Root Cause Analysis. Tracing a failed deployment back to its artifact, build, commit, and any policy or infra changes normally depends on knowing which systems to query, tribal knowledge held by a few senior engineers. Following deployment → artifact → build → commit → recent changes automatically surfaces the likely cause. Enables root-cause tracing that doesn't depend on who's on call.

Cost Attribution. Attributing cloud spend to a team requires joining billing, service ownership, and deployment activity — three systems with no shared key. Joining cost → namespace → ownership → recent deployments tied to specific changes. Enables cost conversations grounded in what changed, not just what was billed.

What We Learned

Five principles emerged from building this.

Start with relationships, not entities: Comprehensive entity modeling is tempting. Resist. Identify three high-value relationships. Model end-to-end. Prove value. Then expand. A build → artifact → deployment path working correctly teaches you more than 50 entity types modeled but unconnected.

Schema as code is non-negotiable: Version-controlled, code reviewed, automatically validated. Not documentation that drifts. Executable source of truth. Your query compiler must read this at runtime.

Data quality is an operational discipline: Track completeness, correctness, and freshness per entity type. Dashboard it. Alert on degradation. Each domain team owns its quality metrics. The graph is only as trustworthy as its data, and incomplete data produces incomplete reasoning.

Governance is workflow, not gate: Don't create approval boards. Create automated contracts: tests that must pass, evaluations that maintain accuracy, and monitors that stay green. If schema changes break tests, they don't deploy. Simple.

Continuous evaluation from day one. Build the harness when you build the graph, not after. Run on every change. Graphs that work today can degrade silently tomorrow. Automated evaluation catches it before users do.

Instrument everything for observability. Evaluation harnesses test synthetic questions. Real users ask questions you didn't anticipate. Trace every query: what was asked, what HQL was generated, which relationships traversed, and what failed. Failed queries expose gaps in relationships and entity coverage. Usage patterns drive schema evolution. Not roadmap planning, observing what breaks in production.

Conclusion

Knowledge graphs don't eliminate complexity—they make it explicit and queryable. The data already exists; the connections don't. Teams that treat schemas like APIs, data quality like operational metrics, and relationships like first-class contracts unlock reasoning across domains. Start with three high-value relationships, prove value end-to-end, then expand. The graph gets smarter with every connection you add.

← Previous:
Next: →‍

Related Resources

No items found.

Get Started

Get Started with Harness AI

Try the full platform free. No module restrictions, no credit card.

Akshay Khole
Principal Software Engineer
Akshay is a Principal Software Engineer focuses on building production-grade knowledge graph at scale
akshay-khole
Akshay Khole
Arvind Aravamudhan
Senior Director, Software Engineering
Arvind is a Senior Director, Software Engineering, who leads the Harness platform engineering team.
eng-lead-arvind-aravamudhan
Arvind Aravamudhan