What Should AI Be Looking for When It Watches a Teacher?
When I first read about Multiverse using AI to evaluate its instructors, part of the idea made sense to me.
One of the problems with teacher evaluation is that an administrator or coach usually sees only a tiny portion of what a teacher actually does. A formal observation might capture thirty minutes out of weeks of teaching, and the teacher knows that someone is watching.
AI seems to offer a solution. If a system can review many lessons instead of one or two, it can potentially identify patterns that a human observer would never see.
More observation should mean better understanding.
But does it?
The more I thought about this through Caddie Alford’s discussion of doxa and bodies in Entitled Opinions: Doxa After Digitality, the less certain I became.
More evidence of what a teacher does does not necessarily give us better access to why the teacher did it.
And that difference matters if our goal is not just to evaluate teachers, but to help them become better teachers.
When Teaching Becomes Data
Multiverse developed an AI system called Compass that reviews its instructors’ teaching sessions against a framework for effective instruction.
The company describes Compass as a way to move beyond occasional observations. The system can review every session, identify possible areas of concern, combine that information with learner feedback, and suggest areas for development. Human managers still review the information and make decisions.
So this is not simply a story about an AI replacing administrators.
But some instructors have described the experience very differently. In reporting by The Guardian, teachers talked about feeling continuously monitored and worrying that ordinary features of teaching—a pause, a filler word, a self-correction, an adjustment made in response to a learner—could become evidence of poor performance.
That initially made me wonder whether AI could evaluate teaching accurately.
I now think there is a more interesting question:
Does having more evidence of observable teaching behavior necessarily mean having better access to the professional judgment producing that behavior?
“Doxa Predispose Bodies”
Alford writes that “Doxa predispose bodies.”
I find that phrase especially useful for thinking about teachers.
She is not describing doxa simply as opinions that we consciously hold and then decide to act upon. Our accumulated beliefs, habits, expectations, and experiences also shape how we encounter the world in the first place.
They influence what catches our attention.
They shape what feels important.
They help determine which responses seem possible.
Teaching works this way constantly.
A student looks away.
Another stops participating.
Someone begins calling out.
A child who was engaged thirty seconds ago suddenly puts their head down.
The teacher usually does not stop everything, gather the available evidence, compare several possible interventions, and then consciously choose the highest-scoring response.
Something catches the teacher’s attention, and the teacher responds.
Maybe I use the student’s name.
Maybe I move closer.
Maybe I ask a question.
Maybe I change the pace of the activity.
Or maybe I notice the disengagement and intentionally do nothing because redirecting one child at that moment might disrupt the rest of the group.
Those decisions can happen very quickly.
Experience shapes what a teacher notices and what kinds of responses seem appropriate. Professional judgment, then, is not only something a teacher thinks. It is also something a teacher enacts.
But Professional Judgment Is Not Automatically Right
This does not mean teachers should simply trust their instincts.
Alford also challenges the idea that we ever look at the world with an “innocent eye.” What we notice is shaped by habits, assumptions, expectations, and ideologies.
That matters in teaching too.
A teacher can notice the wrong thing.
We can interpret a student unfairly.
We can repeat an ineffective habit because it has become familiar.
We can overlook an alternative that another teacher or coach might see immediately.
Experience can produce expertise, but experience can also reinforce assumptions.
For me, this makes reflection more important, not less.
If professional judgment develops partly through experience, then teachers need opportunities to slow that judgment down after the fact and look at it more carefully.
Why did I notice that student?
Why did I interpret the behavior as disengagement?
Why did I intervene immediately in one situation but wait in another?
What other response could I have made?
Would I make the same decision again?
More Visibility Does Not Mean Complete Understanding
Later in the chapter, Alford writes that “no body is ever thoroughly articulated.”
That idea helped me see the limitation of an AI evaluation model more clearly.
We can collect more data about a teacher without completely understanding the teacher.
Suppose an AI system identifies a long pause during instruction.
It can measure the pause.
It can compare that pause with others.
It might even identify patterns between pauses and student outcomes.
But what does the pause mean?
Maybe the teacher lost their train of thought.
Maybe the teacher was deliberately giving a learner more processing time.
Maybe the teacher noticed confusion and was deciding how to explain something differently.
The visible behavior can look almost identical while the professional judgment behind it is completely different.
The same is true of a teacher who appears not to respond.
A student may be disengaging, and the teacher may deliberately choose not to redirect them because doing so would interrupt the activity or draw even more attention to the behavior.
That judgment might still be wrong.
But we cannot know whether it was a good or bad judgment simply by observing that no redirection occurred.
The More Important Question Is What Happens Next
This is where my thinking about Compass changed.
My main concern is no longer that AI might get the score wrong.
Human observers get things wrong too.
An administrator can watch a teacher, misunderstand what happened, and assign a score without understanding the reasoning behind the moment.
AI may actually be very useful for identifying patterns that no administrator could realistically observe.
The more important question is:
What happens after AI identifies something?
One possibility is evaluative.
AI identifies a behavior.
A manager reviews it.
The behavior contributes to a judgment about the teacher’s performance.
Even when a human makes the final decision, the teacher largely remains the object being analyzed.
But there is another possibility.
The same AI-identified moment could become the beginning of a conversation.
What did you notice here?
What were you trying to accomplish?
What alternatives did you consider?
Why did this response make sense at the time?
Looking at it now, would you make the same decision?
Those questions seem much more interesting to me educationally.
They do not assume that the AI has discovered the correct interpretation.
They also do not assume that the teacher’s explanation must be correct.
The teacher brings knowledge of what they perceived and intended. A coach or administrator brings another professional perspective. The coach can challenge assumptions, suggest alternatives, or disagree with the teacher’s reasoning.
The purpose is not to prove that the teacher was right.
The purpose is to make professional judgment available for examination and revision.
What Is Teacher Supervision For?
This realization also changed how I was thinking about the role of administrators, instructional coaches, and teacher educators.
Their responsibility cannot only be to determine whether teachers are performing effectively.
A constructive goal should also be to help teachers develop their capacity for professional judgment.
A teacher who receives a score and corrects one visible behavior may know how to behave differently the next time a nearly identical situation occurs.
But classrooms rarely reproduce identical situations.
Professional judgment matters precisely because teachers constantly encounter situations for which there is no simple rule.
A coach cannot prepare a teacher for every future classroom moment by supplying the correct answer to every past one.
What a coach can help develop is the teacher’s ability to:
notice more carefully,
consider alternatives,
explain a decision,
question that explanation,
and make increasingly informed judgments in situations neither the teacher nor the coach has encountered before.
Two Very Different Roles for AI
Seen this way, AI could enter teacher development in two very different ways.
In an evaluative model, AI expands the administrator’s capacity to observe.
It detects patterns, classifies behaviors, and gives managers more information with which to judge performance.
Human oversight may make that process fairer, but the central purpose remains evaluation.
In a developmental model, AI could help identify moments worth revisiting.
The teacher returns to the moment.
The teacher explains what they noticed and why they responded as they did.
A coach or colleague probes that explanation.
Alternatives are considered.
Assumptions are challenged.
The teacher returns to practice with a slightly different way of seeing.
AI has not made the professional judgment.
It has helped create an opportunity for the teacher to develop it.
A Better Goal Than a Better Score
I do not think the question is simply whether AI should or should not be used to evaluate teachers.
And I do not think professional judgment should become a phrase we use to defend whatever a teacher happens to do.
Professional judgment needs to be challenged.
It needs to be compared with alternatives.
It needs to be revised.
But if the goal is genuinely teacher growth, then the endpoint of supervision should not simply be a more precise score.
It should be a teacher who becomes increasingly capable of noticing what matters, explaining why a response seemed appropriate, questioning that reasoning, and making better judgments the next time the classroom presents a situation that cannot be reduced to a rule.
Perhaps the most useful role for AI in teacher development is not to understand teachers for us.
It may be to help teachers understand their own practice more deeply.
Sources / Further Reading
- Alford, C. (2024). Entitled Opinions: Doxa After Digitality. University of Alabama Press.
- Booth, R. (2026, September 28). “Teachers at Euan Blair firm report ‘horrendous stress’ after AI used to rate their work.” The Guardian.
- Team Multiverse. (2026, September 25). “How We’re Using AI to Invest in Coach Development.” Multiverse.
Leave a comment