She taught high-school English for nineteen years and resigned when her district mandated an AI essay-grading platform, because what she was arguing — and what her students' petition named — was that grading an essay and reading a student are not the same act.
The Essay as Conversation
For nineteen years, Sandra Mok taught eleventh-grade English at Roosevelt High School in Sacramento. She assigned three major essays per semester — personal narrative, persuasive argument, literary analysis — and she read every one of them. Not just for grammar and structure, but for voice. For the thing underneath the argument. For the student.
She had read essays by kids who were falling apart and didn't know how to say so. She had read essays by kids who were extraordinary writers and didn't know that either. She had read the 3 a.m. desperation of a student who cared too much and the competent blankness of one who had learned to give her what she thought she wanted. She graded all of them, but grading was, she would tell you, a side effect. The main event was the conversation the essay opened.
In the fall of 2023, the Sacramento school district piloted an AI essay-grading platform across fifteen high schools. Sandra's school was selected. Teachers were told the tool would "reduce grading burden and provide students with faster, more consistent feedback." Sandra attended the training session, used the platform for four weeks, and submitted her resignation in November. Her students, eleven of whom organized a petition to the school board, cited the same concern she did: that something essential to the relationship between reading and being read had been removed.
The Promise and Limits of Automated Assessment
AI essay-grading platforms — tools that use natural language processing to evaluate student writing on dimensions including structure, coherence, vocabulary, and grammar — have been commercially available for more than a decade. The current generation of large-language-model-based tools represents a significant leap in capability: they can produce rubric-aligned feedback that is more specific and, in controlled studies, more consistent than average human grader ratings.
Consistency is a genuine value. Human grader reliability in essay assessment is, the research literature confirms, imperfect. A 2023 meta-analysis of writing assessment studies found substantial inter-rater variability — the same essay scored by different teachers can vary by a letter grade or more, depending on the grader's background, fatigue, and implicit biases ([Educational Psychology Review, "AI and Human Essay Assessment," 2023](https://link.springer.com/journal/10648)). AI grading systems eliminate most of this variability. They apply the rubric the same way every time.
But consistency is not the same as validity. The question of what essay grading is actually measuring — and whether what it is measuring is what matters for student development — is separate from the question of whether the measurement is consistent. A tool that consistently measures the wrong thing is not an improvement over a human who inconsistently measures the right one.
The research on what automated essay scoring systems measure and what they miss is extensive and somewhat unsettling. A 2023 study in the Journal of Educational Measurement found that AI scoring systems showed significantly higher agreement with human graders on essays that scored in the middle of the rubric than on essays at either extreme — precisely the essays where the human judgment most matters ([JEM, "Automated Essay Scoring and Outlier Detection," 2023](https://onlinelibrary.wiley.com/journal/17453984)). Essays that were unconventionally excellent — strong voice, original argument, unexpected structure — scored lower on AI systems than on experienced human raters. Essays that were competently written but intellectually thin scored higher.
What Teachers Actually Know
Sandra Mok's resignation letter, which she shared with me, did not argue that AI grading was wrong on the rubric. It argued that the rubric was not the point.
"When I read an essay," she wrote, "I am reading a student. The essay is a record of what they were thinking on the day they wrote it, which is never entirely about the prompt. When an algorithm reads an essay, it reads the text. These are different acts. The district is substituting one for the other and calling it efficiency."
This is not anti-technology romanticism. It is a pedagogically defensible position about what assessment is for. The National Council of Teachers of English has issued policy statements arguing that automated essay scoring, however technically capable, cannot fulfill the formative function of teacher-read writing assessment — the function that identifies student development, surfaces individual learning needs, and makes the student visible as a person rather than a score ([NCTE Position Statement on Automated Writing Evaluation, 2023](https://ncte.org/statement/automated-writing-evaluation/)). Sandra's argument echoed these statements, but it came from nineteen years of practice rather than from policy.
Ruha Benjamin has written about what she terms "the default of legibility" in automated assessment systems — the tendency to reward what can be measured and to effectively penalize what cannot, in ways that may encode existing inequities ([Ruha Benjamin, Viral Justice, 2022](https://press.princeton.edu/books/hardcover/9780691222882/viral-justice)). For essay grading, the legibility problem is specific: AI systems trained on rubric-aligned examples tend to reward conformity to expected genre conventions and penalize unconventional writing, including the unconventional writing that often characterizes the most intellectually original students and the students who are still finding their voice.
The Student Petition
The petition that eleven Roosevelt High School students submitted to the Sacramento school board in December 2023 was itself, Sandra noted, an excellent piece of persuasive writing. It cited research on automated assessment reliability. It quoted Paulo Freire on the dialogic nature of education. It argued, with more precision than the district's own implementation documentation, that the pilot had substituted data processing for dialogue without adequate consideration of the consequences.
The board heard the petition. The pilot continued.
"They listened," said one of the petition's student authors, a seventeen-year-old named Joaquin. "They just didn't change anything. Which I guess is its own kind of lesson about institutions."
What the students named, and what the research on AI grading is beginning to document more formally, is a concern about what assessment signals to students about what is valued. When an AI grades your essay, the implicit message is that the relevant properties of your writing are the ones that can be measured by an algorithm — structure, vocabulary, argumentation patterns, grammatical correctness. The properties that cannot be measured — originality, intellectual courage, the specific character of your developing mind — are invisible to the system. Over time, the signal shapes what students produce.
A 2024 report from Stanford HAI examined AI grading deployment across twenty-two school districts and found that in twelve of them, teachers reported observing increased stylistic homogeneity in student writing within two to three semesters of AI grading implementation ([Stanford HAI, "AI in K-12 Assessment," 2024](https://hai.stanford.edu/research/ai-index-2024)). Students appeared to be optimizing for the system — producing essays that scored well — in ways that reduced the range and originality of their writing. This finding was preliminary and based on teacher observation rather than controlled study, but its direction is consistent with the theoretical predictions.
Who Benefits, Who Pays
AI essay grading tools generate real efficiency gains. They dramatically reduce the time teachers spend on grading — a significant source of teacher burnout and overwork, particularly in districts with large class sizes. They produce faster feedback to students, which has documented benefits for learning when the feedback is timely and specific. They reduce the variability in grading that can disadvantage students who have unlucky grader assignments.
Districts that deploy them also reduce labor costs: if teachers spend less time grading, the same teacher can manage more students, or the same budget can stretch further. This is not a secret agenda; it is an explicit part of the ROI case that vendors make to districts.
The costs fall on students who may be developing their writing in a system that cannot see what's most important about it, on teachers whose professional judgment is being replaced by algorithmic consistency, and on the longer-run institutional capacity to identify and develop the students — often from non-dominant linguistic or cultural backgrounds — whose writing is unconventionally strong in ways that standardized rubrics are poorly designed to recognize.
The BLS projects significant growth in demand for high-order writing and communication skills through 2033, precisely as AI grading becomes common in the educational systems that are supposed to develop those skills ([BLS Occupational Outlook Handbook, "Fastest-Growing Occupations," 2024](https://www.bls.gov/ooh/fastest-growing.htm)). The interaction of these two trends is worth examining seriously.
What This Means for You
For teachers and educators: Your professional judgment about what constitutes meaningful writing assessment is backed by substantive research, not merely tradition. If your district is moving toward AI essay grading, request access to the validity data for the tool — not just inter-rater reliability data, but evidence that the tool's scores correlate with longer-term student writing development outcomes. Engage your union or professional association if implementation is mandatory. Document, concretely, cases where AI scores diverge from your own assessment in ways that you can articulate pedagogically.
For school administrators and curriculum directors: The efficiency case for AI essay grading is real. So is the validity case against it as a wholesale replacement for teacher-read assessment. A model that uses AI grading for low-stakes, high-volume formative feedback and reserves teacher reading for meaningful assessments may capture efficiency gains without eliminating the teacher-student dialogue that defines formative assessment at its best. Treat the technology as a tool within a pedagogical framework, not as a replacement for the framework.
For school boards and policy-makers: Before mandating AI essay grading tools in public schools, require independent validation of whether the tool's scores correlate with long-term student writing development, disaggregated by demographic group and first-language background. This research exists in small-scale form; it should be conducted at scale, with the specific tools being deployed, before deployment is mandated.
Sandra Mok is substitute teaching now, on her own terms, in Sacramento-area schools. She grades her students' writing by hand. She is looking for what she's always looking for: the thing underneath the argument. "You can't grade a person," she told me. "You can only read them. If you stop reading them, you stop teaching them. Everything else is bookkeeping."




Comments (0)
Join the conversation!