OpenAI introduced a new way to connect ChatGPT Work and Codex activity to business value on September 16, 2026. Analytics in the Admin Console can now bring usage and cost data, classified task insights, and available outcome metrics into one place so administrators can see where AI is being used, what work it supports, and where a rollout may need training, tighter controls, or more investment.
The important change is not another dashboard. It is the shift from asking “How many people used AI?” to asking “Which workflows changed, what outcome improved, and was that improvement worth the cost and risk?”
OpenAI’s official video demonstrates how an administrator can review team adoption, investigate use cases, inspect Codex contributions, and use the Admin plugin to create a leadership report. This guide turns that demonstration into a practical measurement framework for organizations.
Source note: This article reflects OpenAI’s official announcement, video, and help documentation available on September 16, 2026. Product availability, plan requirements, metrics, classifications, APIs, and administrator controls may change. The rollout framework and governance recommendations below are independent implementation guidance, not undisclosed OpenAI product features.
ChatGPT Work analytics at a glance
| Analytics layer | What it can show | The decision it should support |
|---|---|---|
| Adoption | Active users and usage trends | Where enablement or access may need attention |
| Spend | Credits and token usage | Whether costs match priority workflows |
| Work patterns | Use cases and specific task categories | What people are actually doing with AI |
| Configuration | Model, reasoning, speed, plugins, and skills | Whether the workflow is using the right setup |
| Engineering outcomes | Codex contributions to merged commits, lines of code, and code review | Where to investigate delivery impact with engineering leaders |
| Business outcomes | Team-owned measures such as cycle time, quality, revenue, or capacity | Whether the workflow creates measurable value |
These layers are related, but they are not interchangeable. High usage is not proof of value. Low usage is not automatically a failed rollout. A team may use AI heavily on low-priority work, while a small specialized team may create meaningful value with limited but well-designed use.
What is new in the Admin Console?
OpenAI says the analytics experience brings together three questions that organizations usually investigate separately.
1. Who is using ChatGPT Work and Codex?
The Usage view can show active users, credits, and token activity across ChatGPT Work and Codex. Administrators can filter by group or user to identify where adoption is growing, where spend is concentrated, or where a team has access but little activity.
This is useful for directing support. If a group has low adoption, the next step should not be a generic “use AI more” message. The administrator and team owner should review whether the team has a valuable starting workflow, the correct access, relevant data connections, enough training, and clear safety boundaries.
2. What work are teams doing with AI?
The Insights view uses a task classifier to group a sample of messages into predefined use cases and specific tasks. OpenAI’s example includes software engineering work such as feature development and code maintenance, and sales work such as account research and planning.
The overview reveals the mix of work. A detailed use-case table can show associated credits, messages, and active users. Group filters make it possible to compare how different teams use the tools.
Task classification creates a better conversation than raw activity counts, but it remains an analytical signal. Administrators should validate the categories with team owners and avoid using inferred task labels as employee-performance judgments.
3. What outcomes are connected to the activity?
For engineering, OpenAI’s Outcomes view can show Codex contributions to merged commits and lines of code, together with code-review activity. Filters can help administrators and engineering leaders examine changes over time, by group, user, or repository where available.
This is more useful than counting prompts, but it is still not a complete productivity measure. More contributed code can mean useful capacity, unnecessary code, greater review burden, or a change in repository mix. Engineering leaders should compare these signals with review time, escaped defects, rework, incident rates, maintainability, and delivery lead time.
Why usage is not the same as ROI
An AI dashboard can show what happened inside the platform. ROI requires evidence that something valuable changed outside the platform.
A credible measurement chain looks like this:
- Access: the right team can use the approved AI tools;
- Adoption: people use them in real workflows;
- Behaviour: specific tasks or processes change;
- Operational outcome: time, quality, throughput, or customer experience improves;
- Financial outcome: the improvement creates capacity, reduces cost, protects revenue, or supports growth;
- Risk adjustment: review, errors, compliance work, failures, and change-management costs are included.
Stopping at step two creates an adoption report, not an ROI report.
For example, a sales team may use more credits for account research. The meaningful questions are whether preparation time fell, research quality improved, representatives entered conversations with better context, and those improvements affected pipeline movement or win rate. The AI spend is only one input in that calculation.
How to interpret Codex contribution metrics responsibly
Codex contribution data can help engineering leaders investigate where AI-assisted development is becoming part of delivery. It should not become a simplistic leaderboard.
Use the metrics to ask:
- Which repositories and types of work show sustained Codex contribution?
- Does review time rise or fall as AI-assisted contributions increase?
- Are developers shipping comparable work with fewer interruptions or less rework?
- Do defect rates, incidents, security findings, or rollback rates change?
- Which skills, plugins, models, and reasoning settings support successful work?
- Which teams need better specifications, tests, repository context, or training?
Avoid interpreting accepted lines of code as value by itself. A small, safe deletion may be more valuable than a large generated feature. The goal is better software outcomes, not maximum generated volume.
Where the Admin plugin fits
The Admin plugin brings supported analytics and administrative capabilities into a ChatGPT Work or Codex conversation. OpenAI’s demonstration shows an administrator asking for a 30-day view of Work and Codex usage, credits, team use cases, task breakdowns, and engineering contributions, then requesting a short leadership deck with charts, takeaways, and next steps.
That can reduce manual reporting work. It also introduces an important control boundary: reading analytics is different from changing limits, permissions, groups, or access.
OpenAI describes the plugin as permission-aware, meaning it operates within the user’s existing role and workspace controls. Organizations should still separate read-only analysis from write actions, require review for broad changes, preserve audit records, and verify downstream results before reporting that an administrative action completed.
A practical monthly AI rollout review
A useful monthly review can fit on one page or a short leadership deck if it answers the right questions.
Adoption and spend
- active users by team;
- adoption trend over the previous 30 days;
- credits and token usage by team and workflow;
- high-growth or unexpectedly low-usage groups;
- material changes in model, reasoning, or speed choices.
Work and enablement
- top use cases and task categories;
- plugins and skills supporting those tasks;
- workflows with repeated success;
- teams missing an expected connector, skill, or training path;
- tasks that should not use AI or require additional approval.
Outcomes
- one team-owned outcome for each priority workflow;
- comparison with the pre-AI baseline;
- quality, safety, and reviewer-effort measures;
- Codex contribution trends with engineering-quality indicators;
- areas where the data is incomplete or cannot support a conclusion.
Decisions
- continue, improve, expand, restrict, or stop each pilot;
- identify the owner and next review date;
- fund the workflows with demonstrated value;
- fix training, data access, or governance gaps before buying more capacity.
A 30-day measurement plan
Week 1: choose one workflow and baseline it
Select a task tied to a real business priority. Record its current cycle time, volume, quality checks, failure rate, and people involved. Agree on what a useful improvement would look like before introducing more access or spend.
Week 2: enable the smallest useful group
Give a controlled team the approved model, plugin, skill, and data access required for the workflow. Document permissions, prohibited actions, human approvals, and escalation paths. Train the team on the exact process rather than on generic prompting.
Week 3: review activity and friction
Use Admin Console analytics to check adoption, task mix, spend, and configuration. Interview the team to understand where the dashboard does not capture the full workflow. Fix missing context, unclear instructions, duplicate work, or excessive model settings.
Week 4: measure outcomes and decide
Compare the operational outcome with the baseline. Include review time, correction work, incidents, and employee feedback. Expand only when the result is repeatable and the control model can scale.
Privacy, governance, and measurement limits
OpenAI states that it does not train models on an organization’s business data by default. That policy is important, but it is not the organization’s entire privacy program.
Teams should also determine:
- which conversations and task samples contribute to aggregate insights;
- who can view workspace, group, user, or repository-level analytics;
- how long analytics and exported reports are retained;
- whether employee notice, consultation, or policy updates are required;
- how inferred task categories are validated and challenged;
- which data must never enter AI workflows;
- how leadership reports avoid exposing unnecessary individual details.
OpenAI’s workspace documentation describes analytics as directional organizational signals. Treat them as evidence for investigation, not unquestionable truth. Reporting delays, classification errors, missing downstream outcomes, and incomplete cost attribution can all affect the conclusion.
How I can help measure and improve an AI rollout
I provide AI consulting and custom AI development for organizations selecting high-value AI workflows, designing adoption and outcome metrics, integrating approved business systems, and building governance around ChatGPT, Codex, and other AI platforms.
I can also implement the surrounding system through workflow automation, SaaS product engineering, website development, and mobile app development. The XReporter operations and reporting system demonstrates how structured operational data can become useful reports and accountable workflows.
Book a free strategy call to design a focused AI pilot with a clear baseline, measurable outcome, responsible controls, and an expansion decision your leadership team can defend.
Official sources
- OpenAI: How to connect AI usage to business value
- OpenAI on YouTube: Assess Usage and Value of ChatGPT Work
- OpenAI: Introducing the Admin plugin for ChatGPT Work and Codex
- OpenAI Help Center: Workspace analytics for ChatGPT Enterprise and Edu
- OpenAI Help Center: Reviewing Work and Codex usage and Personal Analytics
Frequently asked questions
What does ChatGPT Work analytics measure?
OpenAI says the Admin Console brings together active users, credits, token usage, task and use-case insights, and available outcome metrics across ChatGPT Work and Codex. Administrators can filter views by group or user to investigate adoption and spending patterns.
Can ChatGPT Work analytics prove AI ROI?
Not by itself. Usage, credits, task categories, and Codex contributions are evidence about activity and adoption. A credible ROI analysis must connect them to independently measured outcomes such as cycle time, quality, revenue, cost, customer satisfaction, or reduced rework.
What Codex outcomes can administrators review?
OpenAI's September 2026 material says the Outcomes view can show Codex contributions to merged commits and lines of code, alongside code-review activity. These indicators should be interpreted with quality, defect, review-time, and rework data rather than used as stand-alone productivity scores.
What is the ChatGPT Work Admin plugin?
The Admin plugin is a permission-aware interface that lets authorized administrators explore adoption, spend, tasks, and supported workspace operations in a ChatGPT Work or Codex conversation. OpenAI says it can also turn analytics into reports and leadership materials.
How should a company begin measuring AI adoption?
Choose one important workflow, establish a baseline, identify the team and approved tools, define one operational outcome and one risk metric, run a controlled pilot, and review the results with the business owner before expanding access or budget.
