Skip to content

Sense AI · Operations and managed improvement

Operate the workflow outcome, not just the model endpoint.

We instrument identity, evidence, policy, model, tools, approvals, workflow state, and user-visible outcome in one service view. Incidents can be contained, explained, reconciled, and converted into controlled improvement.

Service
Workflow outcome, users, states, dependencies, control obligations, and owners
Signals
Trace, task quality, retrieval, policy, tool receipt, backlog, intervention, and cost
Response
Impact, containment, case preservation, recovery, reconciliation, review, and change
02

01 · Service model

A model endpoint can be healthy while the business workflow is failing.

We define the service boundary around user-visible outcome, workflow states, dependencies, controls, and ownership.

01

Outcome and SLI

Completion, acceptance, time, quality, backlog, review effort, correction, and downstream outcome by case class.

02

Dependency map

Identity, source, index, policy, model, tool, approval, queue, consumer, and support dependency.

03

Ownership map

Service, product, domain, source, platform, security, model, integration, incident, and change owner.

03

02 · Observability

One trace should connect request, evidence, decisions, actions, and outcome.

Events are designed for explanation, alerting, reconciliation, and evaluation while following explicit minimisation and retention rules.

01

Context

Request and correlation ID, tenant, pseudonymous identity, workflow revision, case class, and declared purpose.

02

Evidence and decision

Source candidates, access result, context selected, policy result, model and prompt revision, and evaluation flags.

03

State and action

Transition, timer, approval, tool request and receipt, retry, intervention, compensation, and terminal outcome.

04

Service effect

User-visible result, acceptance, correction, delay, backlog, cost, incident link, and downstream reconciliation.

04

03 · Response

AI incidents include wrong outcomes, broken boundaries, stuck work, and silent dependency failure.

Runbooks and exercises distinguish quality, access, data, model, integration, policy, capacity, workflow-state, and operating-process incidents.

01

Detect and classify

User impact, affected cases, workflow revision, dependency, boundary or quality failure, and materiality.

Alert and triage evidence
02

Contain and preserve

Suspend action, constrain traffic, revert revision, route to human work, and retain affected case evidence.

Containment decision and case set
03

Recover and reconcile

Restore dependency, repair state, validate external receipts, notify owners, and reconcile partial outcomes.

Recovery and reconciliation record
04

Learn and change

Add cases, adjust thresholds or controls, fix runbooks, review ownership, and make a new release decision.

Problem record and controlled improvement
05

04 · Technical view

The service view should lead an operator from impact to safe action.

A timeline combines workflow state, dependency spans, policy and approval events, intervention, compensation, and reconciliation instead of collecting unrelated infrastructure charts.

Illustrative reference patternIncident trace and recovery
Workflow incident traceImpact: one action delayed
Request acceptedIdentityrecorded
Source retrievalKnowledgecompleted
Policy decisionControlallowed
Action connectorIntegrationtimeout
CompensationWorkflowcompleted
Contained automaticallyOperator validates receiptCase reconciled
Illustrative pattern. The service trace links technical dependency failure to workflow impact, automatic containment, accountable recovery, and final reconciliation.

Impact first

Operators see affected users, cases, actions, backlog, and downstream exposure before component detail.

Causal trace

The failed dependency and resulting state transition remain connected to automatic or human intervention.

Closed by reconciliation

Recovery is not complete until intended and actual external state agree.

06

05 · Engagement contract

Operations and evaluation meet at every material change boundary.

Handover includes service measures, telemetry, alerts, access, runbooks, exercises, release gates, rollback, case review, and an improvement backlog.

01

Service and telemetry contract

Outcomes, states, dependencies, event schemas, retention, dashboards, and alerts.

02

Response and recovery

Failure taxonomy, triage, containment, suspension, rollback, repair, notification, and reconciliation.

03

Managed improvement

Observed-case review, evaluation updates, threshold and control changes, release evidence, and value tracking.

Engagement contract

What must be true, who owns what, and what leaves the engagement.

A good fit when

  • AI workflows moving into sustained production
  • Services with unclear quality, backlog, or dependency health
  • Teams needing managed monitoring, response, evaluation, and improvement

Required before delivery

  • A named production service and operating owner
  • Access to workflow and dependency telemetry
  • Defined user impact, intervention, and suspension routes

Customer owns

  • Staff incident and domain escalation roles
  • Approve telemetry retention and sensitive-data handling
  • Own service objectives, change acceptance, and residual risk

Delivery outputs

  • Service, dependency, and telemetry model
  • Dashboards, alerts, incident and recovery runbooks
  • Exercises, release gates, case-review and improvement cadence

Not included by default

  • Azure tenant, network, identity, Foundry, Search, monitoring, or platform foundation unless Sense Cloud is included in scope
  • A production guarantee based on a prototype or vendor benchmark
  • Unrestricted autonomous action or implicit approval authority
  • Customer policy, source ownership, or risk acceptance decisions
  • A model-vendor uptime metric presented as end-to-end service health

Questions to answer first

  • Which user-visible failure requires immediate containment?
  • Can every external action be reconciled after interruption?
  • How do incidents and observed cases change the evaluation suite?

The next step is a bounded working session: one workflow, its evidence sources, the decision owner, and the conditions under which the system must stop or hand over.

Assess AI operations