Evaluation
Evaluate AI Code Insights in three phases: prepare, run a two-week pilot, then expand in stages. By the end, you should know whether attribution is credible, deployment is repeatable, and the data is useful.
Define success criteria upfront
- Accurate attribution without workflow friction: Developers can keep using supported agents and their normal Git workflow while DX credibly attributes accepted AI changes in pushed commits.
- Valuable reporting: Teams can see where AI-attributed code is landing, how AI-heavy pull requests move through delivery, and, when session data is enabled, where agent workflows encounter friction.
- Clear improvement opportunities: The reports point teams toward specific changes to enablement, context, steering, task scope, review, or testing that are worth investigating and remeasuring.
Who to involve
- IT: Prepare the managed deployment for Phase 2.
- Engineering leadership: Define success criteria and consume the reports.
- Source control and DX admins: Confirm that the repositories in scope are imported into DX and that the account has access to AI Code Insights.
- Security and privacy: Review endpoint controls, network access, and transcript policy.
Phase 0: Prepare
Confirm access and policy
Before installing the daemon, confirm:
- Your account has access to Admin → AI Code Insights, including installers and API credentials.
- Pilot repositories are imported through a supported source control connector.
- Participants use supported local agents.
- IT and security have reviewed the endpoint and network considerations.
- You have decided whether to enable Transcripts & Agent Experience and who may view session details. See Security and privacy.
Run an end-to-end test
- Install the daemon where the repository and coding agent run.
- Use a supported agent to make a small, recognizable change, make a small manual edit, then commit and push through the developer’s normal workflow.
- Run Verify deployment and confirm the daemon, identity, integration, and repository checks pass.
- Find the commit in AI code percentage and compare the attribution with what the developer expected.
Fix missing or implausible results with Troubleshooting before expanding.
Phase 1: Small-cohort evaluation
Run a two-week pilot with ~15–30 developers from one small team or related repository group. Include the operating systems, agents, security baselines, and remote-development patterns you expect to support.
Brief the cohort
Pilot invitation template: Download the PDF as a starting point for email or Slack. Customize the company name, pilot dates and duration, cohort size, contacts, and transcript policy before sending it.
- Why the organization is evaluating AI Code Insights and what decision it should inform.
- What the daemon collects, what remains on the developer’s machine, and whether session transcripts are enabled.
- That the data will not be used as a proxy for hours worked or individual performance.
- Where participants should report problems.
Ask participants to do three things
- Install and verify: Follow Installation and complete a known-commit check.
- Spot-check attribution: Review representative commits in AI code percentage and compare them with the AI changes accepted during the session.
- Report friction: Share the operating system, agent and version, repository, commit SHA, and a short description of any installation, attribution, or performance problem.
Review the evaluation weekly
Use the same three questions at every check-in:
- Is attribution credible and is normal development unaffected?
- Are the reports producing useful leadership metrics?
- Is the data revealing specific opportunities to investigate?
Track issues in one place with the environment, agent, repository, owner, and next action. Confirm that quiet participants completed their checks.
Prepare managed deployment in parallel
If the evaluation may expand, use Phase 1 to prepare the enterprise deployment:
- Package the installer and managed configuration for your deployment tooling.
- Test on two or three representative machines across the operating systems, hardware tiers, and security baselines in scope.
- Validate endpoint protection, proxies, Git URL rewrites, and agent hooks.
- Decide who owns upgrades, credential rotation, troubleshooting, and removal.
Complete this work before Phase 2.
Exit criteria for Phase 1
Move to a staged rollout only when:
- Installation is repeatable, with no material performance or security issues.
- Known commits look credible across the representative agents and environments tested.
- Blockers are resolved or have a documented owner and plan, including a tested managed deployment path if you plan to use one.
- The evaluation question is answerable, or the missing source, coverage, or sample-size requirement is understood.
Phase 2: Staged rollout
Expand one complete group at a time
Expand one complete team, repository group, or development environment at a time. Keep early waves within validated platforms and deployment paths; test each new environment separately.
Monitor coverage and blockers
For every wave:
- Verify representative machines and at least one known commit.
- Confirm that the expected repositories and agents are reporting.
- Track missing or unhealthy installations separately from report results.
- Compare teams or periods only after coverage is credible.
Use Troubleshooting for common blockers such as endpoint security, proxies, Git URL rewrites, and remote or ephemeral environments.
Pause expansion when the same failure begins appearing across several participants. Resolve the deployment pattern before adding another wave.
Exit criteria for Phase 2
Continue expanding only when:
- Expected installations are complete and coverage gaps are understood.
- Attribution remains credible and performance and security controls remain stable.
- Changes in coverage or cohort composition are documented alongside report comparisons.
Report back
In DX, open Reports and keep the date and cohort filters consistent wherever possible:
- AI code percentage: Validate known commits, record coverage gaps, and find a team, repository, or period worth investigating.
- AI pull request overview: Compare one delivery, review, revert, or work-allocation signal across AI code percentage bands, then inspect representative PRs.
- AI effectiveness: Identify the adoption, delivery, quality-proxy, or session pattern that warrants a closer look.
- Agent Experience Score: If Transcripts & Agent Experience is enabled, choose one friction dimension and inspect the sessions behind it.