Get Ship Done: Everything We Shipped in September 2026

TL;DR
This month's biggest ships
- Harness runs inside ChatGPT now: The Harness and OpenAI plugin lets you manage pipelines, policies, and deployments from ChatGPT on desktop, web, or mobile, using your own Harness access and credentials.
- Merge queues for Code Repositories: Multiple PRs can now be queued and merged safely without manual coordination.
- CD can now import existing AI agents: Agents running on Google Agent Runtime or AWS Agent Core are registered as Agent Services without redeploying, and then managed, promoted, and rolled back through Harness pipelines like anything else.
- Infrastructure as code gets built-in secret scanning: Gitleaks and SonarQube now run as native steps inside infrastructure pipelines, no separate security stage required.
- API discovery and testing are no longer two separate steps: APIs found through code repository scanning now automatically flow straight into API security testing.
- Worker agents can finally reach the model keys that security teams actually hold: New connectors pull Anthropic credentials through your existing AWS, GCP, or Azure cloud provider connector, and an OpenAI-compatible connector now supports LiteLLM as a routing layer.
- The Harness CLI now installs through Homebrew: One command on macOS, no curl pipe. Agent-issued commands already outnumber script-issued ones this month, with Claude Code the top model behind them.
Code is already running at AI speed. Everything after it is still catching up, builds, tests, security scans, approvals, and deploys. That gap is where this list lives. Harness shipped 102 updates in September, close to one every seven hours. Together they move builds, deployments, security testing, API protection, and cost visibility a little closer to the pace code already moves at.
Here's everything that shipped.
Platform updates
Usage Analytics now shows how deeply your teams actually use each module you pay for, not just what you purchased. Module tabs are live for Platform, Deployments, Builds, Security Testing, and Supply Chain, all running on the newer Dashboards engine, so license consumption and actual usage finally sit side by side. Support tickets can now be flagged Urgent directly from the Harness UI for production-impacting outages or critical security incidents, with confirmation tied to the Harness Support SLA. Subscriptions and Billing settings also added a Transaction History section, so admins can review past transactions and payment status without filing a ticket to ask.

The Harness MCP Server moves past read-only access for infrastructure as code. You can now create and update workspaces, variable sets, modules, and providers, and run full infrastructure workflows end to end, directly from an AI tool like Claude or Cursor.
On the Harness AI side, an Expert Agents page now shows what each agent actually does before you decide to use it. Rules management, previously only available inside AI Chat, can now be managed outside it too, governed by the same role-based access as everything else.

The Harness MCP Server itself added features nearly every few days this month, including tools for pulling AI-SRE alert context directly through MCP clients, an autonomous toolset for self-directed development work, lifecycle resources for Vibe Mode, and expanded GitOps and feature-flag tooling, all on top of the IaCM access it already added.
Getting model access into worker agents got easier for teams where the AI keys and the platform team aren't the same people. A new connector option lets worker agents reach Anthropic through your existing AWS, GCP, or Azure cloud provider connector, using IAM roles or OIDC, instead of requiring a separate set of Bedrock keys that security teams are often reluctant to hand out. The OpenAI-compatible connector now also supports LiteLLM as a proxy layer, so requests get normalized, routed, and centrally configured across models and providers instead of being hardcoded to one.

And Harness now runs inside ChatGPT. After an OAuth connection and picking an account, you get the same access that Harness MCP already offers, on your own credentials and permissions, from ChatGPT's desktop app, web, or mobile.
The command line got the same push. The Harness CLI now installs through Homebrew on macOS, no curl pipe or hunting through GitHub release assets, with plugins installing separately through the CLI itself. The Artifact Registry plugin is now production-ready, pushing Terraform modules straight to a registry and bulk-downloading artifacts by regex, alongside a broader round of CLI updates: a full command-line workflow for AI Evals, scriptable AI code review settings, git-backed input sets for pipelines, and full create-read-update-delete for feature flags from the terminal. Agent-issued commands already outnumber script-issued ones this month, with Claude Code the top model driving CLI usage. The CLI is becoming less of a developer convenience and more of the interface agents use to talk to Harness.
Software Delivery Agent: move changes safely to production
Builds and code review
Code Repositories picked up merge queues this month. It sounds small until you've waited on a release train with no queue: multiple approved PRs can now merge in sequence without someone babysitting the order.

Cache Intelligence now caches Node.js and npm dependencies with zero configuration, no custom path setup required, which cuts dependency install time on repeat builds. The Artifactory Upload step adds OIDC authentication, so you can push artifacts to JFrog Artifactory with short-lived tokens instead of a username, password, or anonymous access. Workload Identity now propagates AWS session tags to the assumed role, and the Save Cache steps (S3, GCS, Azure, generic) can skip paths that don't exist rather than failing the entire step, which matters for optional or conditional build output.
The Azure OIDC plugin supports custom subject claims using Harness expressions, and secrets referenced in CI pipelines now record where they're actually used, down to the account, organization, and project level. A handful of smaller fixes also landed, including clearer codebase step logs that print the delegate task ID, which makes tracing a bad Git clone much faster.
Deployments and GitOps
Kubernetes Rolling, Canary, and Blue/Green deployments now have a Pods tab on pipeline executions, showing pod phase, readiness, restarts, and a diagnostic reason, with full status history, events, and logs even after Kubernetes deletes the pod. That's the difference between guessing why a rollout stalled and actually seeing it.

Harness CD can now import an AI agent that's already running on Google Agent Runtime or AWS Agent Core, register it as an Agent Service without redeploying it, and manage, promote, or roll it back through pipelines from that point forward. AI Verify adds fail-fast criteria: metric health sources can stop verification immediately if a metric produces no data at all, and log health sources can stop the moment an unknown log cluster shows up, instead of running out the full analysis window either way.

CloudWatch health sources now support OIDC, exchanging a short-lived token for temporary AWS credentials instead of storing long-lived access keys. ECS Blue/Green deployments can automatically downsize the old service once traffic fully shifts, removing autoscaling policies, deregistering the scalable target, and zeroing out the task count, which eliminates idle container spend without a manual cleanup step.
On the GitOps side, the Update Release Repo step now posts GitHub commit status checks directly to pull requests, and supports pattern-based file path matching, so one service can update multiple files across a release and raise a single PR instead of needing a separate service per file. The GitOps UI adds an App of Apps tab for managing parent-child application relationships. A batch of smaller GitOps updates shipped alongside: automatic cleanup of temporary branches, custom Helm release names, and custom annotations for clusters, JEXL expression support in the GitOps agent, and the agent now bundles the latest ArgoCD release.
Pipeline governance
When a policy evaluation is set to warn and continue, the step, stage, and pipeline now show "Passed with warning" instead of a plain success, with the warning visible on hover, so a soft policy failure doesn't quietly disappear into a green checkmark.

Pipeline templates now support overriding notification rules per pipeline, and insert nesting goes two levels deep, so a stage template dropped into an insert can itself contain another insert step.
On the SaaS side, policies that need network calls (http.send, net.lookup_ip_addr) have to run on your own infrastructure now, not on Harness SaaS, and policy sets in general can be configured to run on your own infrastructure when they need secrets, handle large payloads, or depend on those network builtins.
Smaller additions: a Pipeline Timeout column on the Executions list, and oversized OPA policy outputs now get stripped and replaced with a placeholder so they stop bloating the database, without touching policy status, deny messages, or severity.
Multi-service releases
Release Orchestration activities can now loop over deployment targets, running a separate iteration per target from a single activity definition instead of duplicating activities for every environment or region.
Input set variables can be marked read-only so they can't be overridden at runtime, and OPA policy evaluation now runs on save for release groups, catching governance problems at configuration time instead of at execution. Phase re-run controls let members of a phase's user group re-run or advance to the next phase, and manual activities inside re-runnable phases can no longer be skipped.
Rounding out the month: a banner that flags when GitX webhook sync isn't enabled, a requirement that every phase specify owners and activities, phase inputs fetchable by a human-readable slug, a new Targets panel in the phase editor, fixed row-click navigation in activity looping, and fewer redundant API calls when checking artifact updates on paused phases.
Infrastructure as code
Ansible playbooks now inherit their scope from the Git connector you register them with, so you can register once at the account or organization level and reuse the same playbook across every project instead of registering it per project.
Playbook runs support Tags, Skip tags, and Limit filters to narrow a run to specific tasks, plays, or hosts without touching the underlying playbook. A Remote Static Inventory option, in beta, reads an INI-format inventory file straight from a Git repository instead of requiring hosts to be entered by hand.

Environment variables and OpenTofu and Terraform variables now support Number, Boolean, and JSON types, not just String and Secret. Workspaces can auto-create their default pipelines on creation (limited GA), so Plan, Provision, and Check for Drift are ready without manual setup. Custom plugin images can pre-bake an OpenTofu or Terraform binary, avoiding runtime provisioner downloads in locked-down environments.
The bigger shift: Gitleaks and SonarQube now run as native steps directly inside infrastructure pipelines. Gitleaks scans IaC files for hardcoded secrets like API keys and tokens, SonarQube scans for vulnerabilities and enforces quality gates, and neither requires a separate security stage anymore. Infrastructure code gets the same scanning discipline as application code, in the same pipeline.
Artifact Registry picked up a few improvements worth grouping together: risk occurrences now show how each one was detected, the Risk Details page got a cleaner layout, and the Overview page added an onboarding call to action plus progress metrics for both pipeline and infrastructure scans.
Security Testing Agent: find and fix risks before production
AutoFix now supports Bitbucket Data Center. Teams running self-hosted Bitbucket get the same AI-generated remediation pull requests that Harness Code, GitHub, and GitLab users already have: a finding comes in, AutoFix opens a PR with the fix, and developers review and merge it through their normal process instead of fixing it by hand.

Runtime Protection Agent: protect applications, APIs, and AI in production
Seeing what you actually have
You can now download OpenAPI specs in bulk for your full API inventory, instead of exporting them one at a time for development, documentation, testing, or security work. Asset Inventory now shows Discovery, Testing, and Protection findings on the same page for a given asset, so you can read an asset's full security posture without jumping between modules to piece it together. When an issue fires, Harness now highlights the exact line, payload, or attribute that triggered it, instead of leaving you to scan the whole payload to find what tripped the alert.

The Policies page was rebuilt around an entity-type and filter model with framework mapping, so finding a policy or checking framework coverage requires a search.

Testing catches more before production
Three new out-of-the-box scan policies shipped with documented performance benchmarks, so you can pick a policy based on how you actually scan, whether that's a daily or weekly cadence or a step inside CI/CD. Scan statuses got simpler too: fewer, clearer states in the UI, with a separate Scan Status Reason field carrying the detailed context instead of cramming it into the status itself. For teams using Code Repository scanning to discover their APIs, testing is now fully automated: a discovered API flows straight into API security testing with no manual handoff, and that same discover-and-test path now works directly against GitLab repositories too, not just GitHub and Bitbucket. DAST added four new GitHub Actions, plus scan overrides, more realistic request generation, and dependency-aware testing that accounts for what a service actually calls.

Blocking attacks without over-blocking good traffic
Threat actor scoring picked up three changes that matter together: score contribution per threat category is now configurable, scores decay for actors that are already blocked instead of staying maxed out forever, and detections are skipped entirely for requests that are already allowed. The net effect is a system that keeps blocking real threats without wasting cycles re-litigating traffic it has already cleared.

Cost Management Agent: control and optimize cloud and AI costs
The improvements here are plumbing, not a feature you'd click on, but they're the plumbing that makes every other cost number trustworthy. Cluster event ingestion, hourly billing sync, and Jira status updates on cost recommendations all got more reliable error handling this month. If you've ever had a cost dashboard quietly fall behind without telling you, this is the fix for that.
Other product updates
Harness Feature Management and Experimentation added a Dart Client SDK and OpenFeature Provider. It runs on the Dart VM and on Flutter across Android, iOS, web, and desktop, written entirely in Dart with no platform-specific dependencies, so it runs anywhere Dart or Flutter runs.
Resilience Testing picked up a dense set of improvements this month. The ones worth knowing about: Helm deploy pipeline scans can now pick up services and risks straight from Helm-based CD pipelines, load test worker counts can be estimated from a formula instead of guessed by hand, scan actions now honor Harness roles and resource groups instead of running wide open, Chaos Probe templates can carry their own risk rules, and k6 load tests now report per-endpoint and per-task statistics alongside P50, P95, and P99 latency numbers for Locust and JMeter runs too. A longer tail of smaller fixes shipped alongside: clearer, consolidated risk scan statuses, clickable risk counts that jump straight to a filtered view, severity sorting on the risks list, a reworked summary card on scan details pages, color-coded request method badges on endpoint statistics, and a side panel for inspecting chaos timeline details without losing your place.
The gap this closes
102 updates, thirty days. Writing code fast was never going to be the hard part once agents got good at it. The hard part is everything a change has to survive after it's written: a build that doesn't break, a test that actually catches the regression, a scan that finds the vulnerability before a customer does, a deploy that rolls back cleanly if it doesn't work. Every item on this list moves one of those stages a little closer to the speed code already runs at. That's the race that actually matters right now, not who can generate a pull request fastest, but who can get it safely into production without slowing down to check.



