Pleasanton

PleasantonCA.city — Tri-Valley neighborhood news with hometown heart.

Back to Pleasanton

Is AI Helping Our Children? California Bet on Teaching It. The Research Is Mixed.

California requires AI literacy in the curriculum. But a randomized trial found students using an unguarded chatbot scored 17 percent worse once it was taken away, and a Stanford study found AI detectors misfire on multilingual students.

Reese Hardy

August 20, 20266 min read

Bar chart of results from a randomized trial of GPT-4 in high school mathematics. During practice, students using a standard chatbot scored 48 percent higher than classmates with no AI access and students using a hints-only tutor scored 127 percent higher. On an exam with the tool removed, the standard-chatbot group scored 17 percent lower, while the hints-only group showed no significant change. Source: Bastani et al., PNAS, 2025.
Bar chart of results from a randomized trial of GPT-4 in high school mathematics. During practice, students using a standard chatbot scored 48 percent higher than classmates with no AI access and students using a hints-only tutor scored 127 percent higher. On an exam with the tool removed, the standard-chatbot group scored 17 percent lower, while the hints-only group showed no significant change. Source: Bastani et al., PNAS, 2025.

California decided some time ago that the answer to artificial intelligence in schools was to teach it rather than fight it. The research that has arrived since suggests that instinct was probably right, and that the details will decide whether it works.

Start with the sharpest finding available. Researchers gave roughly a thousand high school students access to GPT-4 during mathematics practice. Practice scores rose 48 percent. Then the tool was removed for an exam, and those same students scored 17 percent worse than classmates who had never used it at all.

The study, published in the Proceedings of the National Academy of Sciences, describes the students using the model as a crutch.

Why that result is worse than it first sounds

The failure was invisible while it was happening.

Practice work improved. Homework looked better. Engagement looked better. Every number a school actually collects pointed upward, and the loss only surfaced on a test where the tool was absent. A district measuring assignment completion through an AI saturated year would record success while measuring the software instead of the student.

The same model, configured differently, did no harm

A third group in the study used the identical GPT-4 model with one change. It had been instructed to offer incremental hints and never hand over the answer.

Those students improved most during practice, by 127 percent, and showed no significant deficit afterward. The technology was not the variable. The instruction given to it was.

What California has actually required

Assembly Bill 2876, signed by Governor Newsom on September 29, 2024, folds AI literacy into the state's core curriculum work. It directs the Instructional Quality Commission to consider incorporating AI literacy into the mathematics, science and history social science frameworks as those are revised, and into the criteria used to evaluate instructional materials.

The framing in the bill covers both halves of the problem: learning about AI, meaning how it works and how it affects society, and learning with it, meaning using it effectively and ethically. The California Department of Education has since published statewide AI guidance for districts, county offices and charter schools, developed with an artificial intelligence working group.

That is a more considered starting point than most states have. It also leaves the hardest question open, because a framework can require that students learn about these tools without specifying whether the tool in front of a child hints or answers.

The Stanford finding that matters most in the Tri-Valley and the South Bay

There is a second research result with unusually direct local relevance, and it concerns enforcement rather than learning.

Stanford researchers, led by Weixin Liang with James Zou and colleagues, tested widely used GPT detectors against essays written by human non-native English speakers. The detectors consistently misclassified that writing as AI generated. More than half of the non-native TOEFL essays were flagged. On essays by native speaking US eighth graders, the same detectors were nearly perfect.

The mechanism is unforgiving: the tools read the plainer sentence construction of a second language writer as machine output. The researchers showed the bias could be removed simply by prompting for more varied phrasing, which means the detectors were penalizing limited linguistic range rather than detecting cheating.

In districts where a large share of families speak a language other than English at home, that is not an abstract fairness concern. It is a prediction about which students get accused. Universities have acted on exactly that: Vanderbilt disabled Turnitin's AI detector in August 2023, citing its reliability, false positives and the disparate impact on international students.

The practical conclusion is uncomfortable but clear. Policing is not available as a strategy, which puts the weight back on how the tools are configured and supervised.

What teachers report

Gallup and the Walton Family Foundation found roughly six in ten teachers used AI during the past school year, with about a third using it weekly. Weekly users report saving an average of 5.9 hours a week, roughly six weeks over a school year.

Only 18 percent reported receiving formal guidance from administrators on how the tools should be used, and about a third reported none at all. California's statewide guidance is aimed squarely at that gap, though guidance issued to districts and guidance reaching an individual classroom are not the same thing.

The case on the other side

A World Bank randomized trial in Benin City, Nigeria, ran first year senior secondary students through six weeks of teacher supervised after school English sessions using Microsoft Copilot. The gain was 0.31 standard deviations across the full assessment, which covered English, AI knowledge and digital skills, and 0.23 on English alone, the study's main outcome.

The researchers report the program beat 80 percent of the interventions in a comparison database of randomized trials in developing countries.

The sessions were structured and teacher supervised, much closer to the hint giving arm of the Turkish study than to independent chatbot use. The benefits were also uneven: the largest effects were among female students and among those with higher initial academic performance, a reminder that a tool can lift averages and widen gaps at the same time.

The limits of the evidence

One subject, one country, four sessions covering about 15 percent of the semester's mathematics curriculum. Math is unusually friendly to a hint based tutor, because answers are determinate and student errors are well catalogued. Nobody has shown the same design works for an essay.

The horizons are short. Every trial here measures weeks. Whether the exam deficit persists over years, or whether students raised alongside these tools develop other capabilities worth having, is simply unstudied.

One thing this research does not do is explain falling test scores. National assessment scores peaked in 2012, and the declines among the lowest performing students set in from there, years before ChatGPT shipped. No honest reading of the timeline pins them on chatbots.

What a parent can actually check

Nothing in this research argues for confiscating the tools, and nothing in it argues for leaving a child alone with them. It argues for one specific question, which requires no technical knowledge to ask:

When your child gets stuck, does the tool produce the answer, or the next step?

In the only controlled comparison available, that distinction accounted for the whole gap between learning and not learning. Right now it is mostly being decided by product designers, not by families or school boards, and it is worth asking a teacher which version is in the classroom.

Sources

  • Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakci, O. and Mariman, R. "Generative AI without guardrails can harm learning: Evidence from high school mathematics." Proceedings of the National Academy of Sciences, 2025.
  • Gallup and the Walton Family Foundation, teacher AI surveys, 2025.
  • AcadeResearch, "Practice Up 48%, Exams Down 17%: Does AI Stop Children From Learning?", August 2026.
  • California Assembly Bill 2876 (2024), pupil instruction: media literacy: artificial intelligence literacy.
  • California Department of Education, artificial intelligence guidance.
  • Liang, W., Yuksekgonul, M., Mao, Y., Wu, E. and Zou, J. "GPT detectors are biased against non-native English writers." Patterns, 2023.
  • World Bank Education Global Department, "From Chalkboards to Chatbots," Nigeria.
Share

Reese Hardy

Reese Hardy writes about community life, schools, public safety, and local events in Pleasanton.

Related Stories

More in California