AI effectiveness

The AI effectiveness report combines provider adoption, PR outcomes, coding-agent session analysis, and model-fit findings. Use it to find friction, understand agent work, assess model choices, and inspect the sessions behind each pattern.

Note: The page can combine independent data sources. The 0–100 AI Effectiveness Score requires DX AI, at least one AI tool provider connection, and source control data. Agent Experience, friction, work-type, and Model fit analysis require AI Code Insights session data and Transcripts & Agent Experience enabled.

AI effectiveness report in DX

When to use AI effectiveness

AI effectiveness answers the question: Which agent-workflow, model-fit, adoption, delivery, or quality pattern should we investigate?

This report helps teams:

  • Locate recurring friction — Find low requirements, steering, or scope ratings and open the sessions behind them.
  • Understand agent work — See the work types and categories represented in captured sessions.
  • Improve model selection — Find work that may suit a lower-cost or more capable model.
  • Inspect session evidence — Drill into the sessions behind a friction area, work type, Model fit finding, or breakdown row.
  • Compare cohorts and components — Break down available signals by team, group, or attribute and separate adoption, velocity, and quality patterns.

Data sources and prerequisites

Section Source Requirement What it can support
Overall, Adoption, Velocity, Quality AI provider and source control connectors DX AI, an enabled AI tool provider connection, and required PR data A high-level map of adoption, throughput, and bug-fix allocation for AI users
Agent experience AI Code Insights session evaluations Daemon installations and Transcripts & Agent Experience enabled Requirements, steering, and scope patterns
Friction areas Agent Experience evaluations Evaluated sessions Counts of sessions rated 3 or lower on a dimension
Agent work Transcript-derived work classification Captured sessions with non-empty messages Work types and categories represented in sessions
Model fit (beta) Session task complexity and model Assessed sessions using a supported model Potential savings or quality opportunities by work type and model
Breakdown table Available score and session signals Columns vary with connected data sources Patterns by team, group, or selected attribute
Session drilldowns Daemon sessions, transcript access, and linked delivery data Availability varies by integration and viewer permission The session, evaluation, transcript, and linked output behind a pattern

Missing sources can make a section unavailable without invalidating the data shown in another section. Check the source requirements before comparing rows or time periods.

How the AI Effectiveness Score is calculated

The 0–100 Overall score is an equal-weighted average:

(Adoption score + Velocity score + Quality score) ÷ 3

If any component is missing, the overall score is N/A.

Adoption

Adoption measures the share of eligible contributors actively using an AI tool, based on provider connector data. DX normalizes adoption to a 0–100 component score.

The current normalization uses these anchor points: 0% adoption maps to 0, 75% maps to 50, and 100% maps to 100. Values between anchors are linearly interpolated.

Velocity

Velocity uses PR TrueThroughput for contributors classified as light, moderate, or heavy AI users. It requires both provider usage and source control data.

The current normalization maps 0 PRs per contributor per week to 0, 2 to 50, 5 to 90, and 10 to 100, with linear interpolation between anchors.

Quality

Quality is a reverse-scored PR defect ratio for contributors classified as light, moderate, or heavy AI users. Defect ratio is the share of PR work allocated to bug fixes; it is a proxy, not a direct measure of defects caused by AI.

The current normalization maps a 0% defect ratio to 100, 10% to 90, 20% to 50, and 30% to 0, with linear interpolation between anchors.

These component scores are DX normalization scales, not industry benchmarks.

Agent experience and friction

The Agent experience section embeds the 1–5 Agent Experience Score. It evaluates eligible sessions on requirements, steering, and scope using a separate post-session model.

Friction areas count sessions rated 3 or lower on one of those dimensions. A raw count depends on session volume, so compare it with the number of evaluated sessions and use a friction rate in custom analysis when cohort sizes differ.

Open a friction area to review representative sessions. Low requirements can point to issue context or task framing; low steering can point to follow-up guidance; low scope can point to task sizing or mid-session drift.

Agent work

Agent work groups transcript-powered sessions by work type and category. A session can receive up to three work types, so work-type counts are not mutually exclusive and should not be added together as a session total.

Use this section to check whether friction clusters around testing, refactoring, feature work, or another task type before changing a team-wide practice.

Model fit (beta)

Note: Model fit is in beta and available to select customers.

Model fit compares the assessed complexity of each session with the model used for the work. It helps you see when a lower-cost model may be enough or when a more capable model may improve results.

The first three highlights show the model and work-type combinations that affect the most sessions. The results below are organized by work type and model. Select a model name to bring its largest opportunities to the top, or select Work type to sort by name.

Each work type and model pairing has one of four results:

  • Well matched - No cost or quality opportunity applies to 25% or more of the assessed sessions.
  • Potential savings - At least one in four assessed sessions used a model more capable than the work required, so a lower-cost model may be enough.
  • Potential quality improvement - At least one in four assessed sessions used a model less capable than the work required, so a more capable model may improve results.
  • Not assessed - The sessions do not have a task complexity rating.

The session count shows how many sessions have both a task complexity rating and a supported model. A session can be assigned up to three work types, so it may appear under more than one work type. Select any highlighted finding or result to see the sessions behind it.

Model fit currently supports:

  • Claude Fable 5
  • Claude Haiku 4.5
  • Claude Opus 4.5
  • Claude Opus 4.6
  • Claude Opus 4.7
  • Claude Opus 4.8
  • Claude Opus 5
  • Claude Sonnet 4
  • Claude Sonnet 4.5
  • Claude Sonnet 4.6
  • Claude Sonnet 5
  • Gemini 2.5 Flash
  • Gemini 3 Flash
  • Gemini 3 Pro Preview
  • Gemini 3.1 Pro
  • Gemini 3.5 Flash
  • Gemini 3.6 Flash
  • GPT-5 mini
  • GPT-5.1
  • GPT-5.2
  • GPT-5.3 Codex
  • GPT-5.4
  • GPT-5.4 mini
  • GPT-5.4 Nano
  • GPT-5.5
  • GPT-5.6 Luna
  • GPT-5.6 Sol
  • GPT-5.6 Terra
  • Grok 4.5
  • Kimi K2.7 Code
  • MAI-Code-1-Flash

Sessions that use other models are omitted from Model fit. Contact your DX account team if you need another model supported.

Session drilldowns

Drilldowns connect each high-level pattern to the sessions behind it. You can open the sessions behind a friction area, work type, category, Model fit finding, or breakdown row.

  • Overview — Session summary, evaluation, tool, contributor, timing, and token information when provided.
  • Output — Repositories, branches, commits, PRs, and deployments associated with the session when available.
  • Transcript — Scrubbed message history for viewers allowed by the transcript access setting.

Session-to-output relationships are not always one-to-one. A session can produce no commit or contribute to multiple outputs, and an output can combine multiple sessions and human work.

Model fit drilldowns are filtered to the selected work type and model. When available, a Model fit label appears beside the session’s task complexity.

Use session drilldowns to separate systemic patterns from one-off sessions. For example, a high count for Poor upfront requirements or context provided can point to issue templates, planning practices, or prompt guidance that need improvement.

Breakdown by team or group

The breakdown table shows available AI effectiveness signals by team, group, or selected attribute. Its columns depend on the connected data sources: session data adds Agent Experience Score, while provider and source control data add Overall score, Trend, Adoption, Velocity, and Quality.

The AI usage attribute group is excluded from the attribute filter because the report already uses it to identify AI users for Velocity and Quality.

Interpret the report safely

  • The overall score is an internal composite, not a causal measure of effectiveness or ROI.
  • Adoption is based on provider-connected eligible users, not daemon code attribution.
  • Velocity compares throughput for users with AI usage; it does not isolate the effect of AI.
  • Quality uses bug-fix allocation. It does not show that AI caused a defect.
  • Agent Experience measures the interaction, not code quality, employee performance, satisfaction, or business impact.
  • Missing data, partial rollout, tool coverage, work mix, and cohort selection can change the result.

Use the report as a map. Move to the component’s source data, inspect representative records, change one practice, and compare the same cohort in the next period.

How AI effectiveness differs from AI impact

Unlike AI impact, AI effectiveness combines provider, PR, and session signals to help you locate an operational question. AI impact uses productivity patterns, time-savings inputs, and financial modeling for a different analysis.

Use AI effectiveness to investigate how adoption, delivery proxies, and agent workflows fit together. Use AI impact when you need the assumptions and outputs in DX’s impact model.