Skip to content
  1. Home
  2. Guides
  3. Guide

QCon Shanghai Preview: How Harness Engineering Constrains SRE Agents

For teams already operating coding agents and agent frameworks, the production question is no longer simply whether an agent can analyze an alert or suggest a change. It is how to carry work into areas such as release management, scaling, configuration, infrastructure adjustment, and fault repair without losing control of production.

That is the focus of Wang Jie’s confirmed QCon Shanghai 2026 session, “理性驾驭 AI 的 SRE 可靠性工程:用 Harness 工程打造 SRE 可控的生产环境”—a discussion of SRE reliability engineering that uses Harness engineering to create a controllable production environment.

The session’s central argument is practical: model capability alone does not make an SRE agent production-ready. The system around the model must constrain what it can do, how changes are reviewed, how people and agents collaborate, and how completed work becomes reusable organizational knowledge.

Why stronger models do not remove the SRE risk

The published preview identifies several gaps in current SRE Agent practices:

  • Many agents remain read-only, stopping at alerts and fault analysis instead of participating in operational work.
  • LLM output is probabilistic, may contain hallucinations, and can be difficult to reproduce. Those characteristics conflict with production requirements for determinism, accuracy, and consistency.
  • A single-user, single-thread conversation cannot carry the coordination required for team operations.
  • Completed tasks often lack a quality decision, leaving useful experience scattered across individuals and sessions instead of becoming an organizational asset.

The proposed response is not to claim that an LLM has become deterministic. Wang describes the goal as using Harness engineering to manage uncertainty instead.

A code-based, GitOps foundation for agent work

The case study is based on Tencent Games’ code-based application delivery and GitOps engine. In this environment, a release change is represented as a change to code or configuration—work that falls within an area the preview describes as comparatively mature for AI coding.

This foundation serves two purposes for agents:

  1. It gives the AI stable context for the change it is being asked to make.
  2. It establishes a guardrail that the agent cannot bypass: changes must pass through Git and human review.

That distinction matters. An agent is not simply granted permission to act on production. It receives an execution path whose required stages remain in place even when the model is uncertain. The engineering harness constrains the route into production rather than relying on the model to behave cautiously on its own.

Human review and multi-agent coordination

The Tencent Games practice described in the preview allows multiple SRE agents and people to collaborate on operational tasks, including:

  • Release changes
  • Scaling
  • Configuration and infrastructure adjustments
  • Fault repair

One planned case involves an SRE Agent Team coordinating the expansion of cloud infrastructure and game applications. The public material does not disclose a specific orchestration protocol or assign fixed roles to individual agents. It does, however, make the coordination problem explicit: production work requires collaboration between human reviewers and multiple AI agents rather than an isolated chat interaction.

Human review is positioned as part of the mandatory change path. Git and review therefore act as production controls, while the agents contribute to execution within that controlled process.

Record, evaluate, then distill

The proposed learning loop has three explicit stages:

  1. Record the complete process of each task.
  2. Evaluate the completed task.
  3. Distill evaluated experience into capabilities that can be shared across businesses.

This is the session’s answer to experiences being trapped inside individual sessions. It also makes evaluation a prerequisite for reuse: a task does not automatically become organizational knowledge merely because an agent reports that it is finished.

What this means for post-mortems

The preview does not describe a conventional incident-postmortem template. It does provide a related operational-learning sequence: preserve the task process, evaluate the result, and turn accepted experience into reusable capability.

For teams running agents in production, useful questions for the session include:

  • Which records are retained for fault analysis?
  • What evidence determines whether a task passes evaluation?
  • How are unsuccessful actions distinguished from reusable successes?
  • Which parts of an incident or repair can safely become shared capabilities?

These questions remain open in the published material; they should not be treated as details the preview already resolves.

Cost and model routing: an explicit gap in the preview

The public preview does not specify model catalogs, routing thresholds, token budgets, price targets, or cost-accounting methods. It therefore should not be presented as a recipe for cost-aware or model-routing architecture.

For teams evaluating the session, cost and routing remain important questions to investigate: which work should use which kind of model, how usage is measured, and what limits apply to repeated or multi-agent execution. The stated contribution is the surrounding Harness design—code and configuration, Git-based changes, human review, collaboration, recording, evaluation, and reuse—not a disclosed model-selection policy.

Questions worth bringing to the session

Teams that already run coding agents or agent frameworks can use the following questions to evaluate the ideas presented:

  • Which SRE tasks are allowed to move beyond read-only analysis into real operational actions?
  • How does the GitOps foundation ensure that changes remain represented as code or configuration?
  • Where are Git and human review enforced in the execution path?
  • How are responsibilities divided and coordinated between people and multiple SRE agents?
  • How is task quality evaluated before an experience is distilled and shared?
  • Which operational records are needed to support review, fault analysis, and future agent improvement?

Why the Harness approach matters

The strongest idea in the preview is not that agents should be given unrestricted autonomy. It is that useful autonomy requires a production system built around enforceable constraints.

For organizations already experimenting with agents, the questions become:

  • Can the agent complete the work?
  • Can the change be represented, reviewed, and traced?
  • Can several agents and people coordinate without relying on an informal conversation?
  • Can the result be evaluated before it becomes reusable knowledge?

QCon Shanghai 2026 will take place on October 22–24. The listed venue is 上海建工浦江皇冠假日酒店.

For teams working toward real agent execution, Wang Jie’s session offers a focused lens on the gap between model capability and dependable operations: design the Harness so agents can act, but keep production control outside the model.

Public sources

  • utm_source=rss&utm_medium=article)
  • QCon Shanghai 2026 schedule
  • QCon Shanghai registration and venue details