Research
From Misconceptions to Evidence: What Science Teachers Make Visible When Co-Designing Agentic Learning Apps
Overview Research area: Human-Computer Interaction, specifically teacher professional learning, educational AI, and co-design of "agentic" learning applications in science education. Technical level:
- arXiv
- 2609.03917
- Published
- 2026-09-03
- Authors
- Nizam Kadir, Wei Ting Liow, Sumbul Khan, Lay Kee Ang
AI summary
Overview
Research area: Human-Computer Interaction, specifically teacher professional learning, educational AI, and co-design of "agentic" learning applications in science education.
Technical level: Beginner-Friendly. The paper is a qualitative, conceptual study with no algorithms, models, or quantitative benchmarks; it analyzes written design artifacts using a five-dimension coding framework.
Scope: A bounded cross-case analysis of four de-identified science-teacher co-design artifacts, asking how the teachers translated disciplinary problems of practice into specifications for AI-supported learning interactions — and which elements of governance they left unstated.
What This Paper Is About
Generative AI now lets anyone sketch an educational app from a short text description, but science teaching is not generic content delivery: it depends on eliciting learners' models, diagnosing misconceptions, interpreting evidence of thinking, and keeping professional judgment in the loop. The problem this paper addresses is that we know little about the small, concrete design artifacts through which teachers translate a science-learning difficulty into an imagined AI interaction. The goal is to examine four such artifacts and see what pedagogical work they made visible — and what they left out.
Key Contributions
-
An empirical description of four science-specific design artifacts. The study analyzes an experimental-design diagnostic (S1), a Kinetic Particle Theory dialogue guide (S2), a chemistry prior-knowledge checker for formulas and equations before mole calculations (S3), and a physics application/scaffolding proposal using real-world scenarios and student video investigations (S4).
-
Identification of an "epistemic specification" problem. All four artifacts articulated a disciplinary problem, a learner interaction, and pedagogically interpretable evidence, but explicit teacher authority appeared in only two and an explicit safeguard in only two.
-
A reframing of app co-design as pedagogical reasoning. The paper argues that teacher professional learning should treat AI app ideation as epistemic specification work rather than feature brainstorming or prompt technique.
-
A five-question design protocol. The protocol — problem, learner interaction, evidence, teacher authority, safeguard — is offered as a conceptual implication to help teachers turn science-learning needs into accountable human–AI arrangements before building or adopting a tool. The paper states explicitly that this protocol has not been evaluated.
Main Findings
-
The cases started from epistemic bottlenecks, not content topics. S1 targeted gaps and misconceptions in experimental design; S2 targeted the difficulty of explaining an invisible particle model; S3 placed prerequisite formulas and equations before mole calculations; S4 targeted transfer of physics knowledge to real-world situations and investigations. None was organized around delivering more information about experiments, particles, moles, or physics.
-
Proposed agency centered elicitation, scaffolding, and repair. S1 used three questions with personalized feedback and reflection; S2 used dialogic questioning and differentiated guidance toward answer formulation; S3 paired a short diagnostic with misconception-sensitive repair; S4 combined scenarios, scaffolded knowledge building, and video investigations. None of the four was centered on producing a finished scientific answer for the learner.
-
Every artifact specified evidence with disciplinary meaning (4/4). S1 would return diagnosed gaps and misconceptions; S2 would make conceptual shortcomings available to the teacher; S3 would produce class-level gaps and the proportion of learners needing support; S4 would preserve investigation performance for teacher evaluation. The envisioned audience shifts from learner to teacher and from individual to class.
-
Teacher authority was explicit in only two of four cases (2/4). S3 gave the teacher authority to select which concepts the diagnostic would assess; S4 reserved evaluation or grading for the teacher and named human review. S1 and S2 did not make teacher authority explicit.
-
Safeguards were explicit in only two of four cases (2/4). S2 stated a privacy safeguard but did not specify direct teacher approval or editing rights; S4 named human review. S1 made neither dimension explicit, and S3 named no safeguard for the resulting misconception data.
-
Authority and safeguards are not interchangeable. The contrast between S2 (privacy named, decision rights not) and S3 (teacher configures content, no data boundary named) shows privacy without decision rights can leave feedback ungoverned, and teacher configuration without a data boundary can leave evidence use underspecified.
-
Coverage counts were descriptive only. The paper states that with four artifacts the counts (4/4, 4/4, 4/4, 2/4, 2/4) are not statistical estimates and make no prevalence claim.
-
No implementation or outcome data. The study explicitly reports no classroom-use, usability, workload, or learning-effect data; the unit of analysis is the content made explicit in each preserved artifact.
Methodology in Plain English
The researchers ran a bounded qualitative cross-case analysis. The material came from a science-teacher professional-learning and co-design activity in which teachers started with a problem of practice and imagined an agentic learning application; one artifact was posted as a group framing and three were later structured proposals. Artifacts were included only if they named a science-specific disciplinary problem and described a responsive application interaction — four met both criteria.
Each case was summarized in a structured matrix across five dimensions drawn from formative assessment, science practices, and teacher-agency literature: disciplinary problem, learner interaction, evidence, teacher authority, and safeguard. A dimension was marked present only when the artifact stated or directly entailed it; the team did not infer teacher control from the educational setting, and did not treat generic data collection as pedagogical evidence. They then wrote within-case summaries and compared cases for common relations and negative contrasts. The authors note they treat each artifact as a design expression rather than an individual participant response, because group reporting and individual authorship cannot be equated. The paper describes "agentic learning application" as a working term for a proposed interaction in which an application delegates a responsive action such as questioning, diagnosis, scaffolding, or feedback — not evidence that an autonomous system was built.
Why This Matters
Impact on research. The paper extends formative assessment logic — evidence becomes formative through use — into the design-specification stage, before software exists. It suggests that shared-control and teacher–AI complementarity research can begin at ideation rather than at implementation, and it offers a coded artifact matrix and a five-question protocol as reusable analytic objects for others to test.
Real-world applications (as proposed orientations in the artifacts, not evaluated products):
- A diagnostic that surfaces experimental-design misconceptions through a small number of questions plus feedback and reflection.
- A dialogue guide that supports learners in formulating and revising explanations about an invisible model such as Kinetic Particle Theory.
- A readiness checker that identifies missing prerequisite formulas and equations before mole calculations and reports class-level gaps plus the proportion of learners needing support.
- A physics scaffolding tool that moves from real-world scenarios to student-recorded video investigations whose performance a teacher evaluates.
Industry relevance. For teams building or procuring educational AI, the paper's practical implication is that specifying "personalized feedback" is not enough: the specification should name who interprets the evidence, at what level of aggregation, for which next action, and which decisions remain professional. The paper frames governance as part of pedagogical design rather than a compliance addendum, which directly affects how products expose teacher controls, data flows, and human-review points.
Future Directions
- Compare specifications produced with and without the five-question protocol, as the paper itself proposes.
- Follow the full provenance chain from problem statement through specification to enactment, examining whether implemented interactions preserve the intended learner work and teacher decision rights.
- Test outcomes empirically. The paper states that whether these designs improve understanding would require implementation, classroom enactment, and outcome measures, including implementation fidelity, technical feasibility, usability, teacher workload, and student experience.
- Widen the analytic frame and invite teacher review of interpretations. The authors note other readings could foreground curriculum alignment, accessibility, assessment validity, or classroom feasibility, and suggest teachers review the analytic interpretations of their artifacts.
Target Audience
Science teacher educators and professional-learning facilitators who run AI or co-design workshops; researchers in educational AI, HCI, and formative assessment interested in specification and governance before implementation; and product, curriculum, or procurement teams in education technology who need to decide what a teacher should be able to select, review, evaluate, or release — and what boundary should prevent unacceptable use. The paper requires no technical background, but readers should note its explicitly bounded claims: it reports what four co-design artifacts made visible, and it reports no implementation or learning-outcome data.
Authors’ abstract
Science educators increasingly encounter AI tools that generate content, yet disciplinary teaching depends on eliciting learners' models, diagnosing misconceptions, interpreting evidence, and preserving professional judgment. This study asks how science teachers translate such epistemic work into specifications for agentic learning applications. It contributes to the conference theme, "Innovating Pedagogies, Inspiring Minds: Transforming Science Learning," and the Teachers' Professional Learning strand by examining app co-design as a form of pedagogical reasoning. We conducted a bounded qualitative cross-case analysis of four de-identified artifacts produced in a teacher professional-learning workshop: an experimental-design diagnostic, a Kinetic Particle Theory dialogue guide, a chemistry prior-knowledge checker, and a physics application/scaffolding tool. Each artifact was coded for the disciplinary problem, learner interaction, evidence made visible, teacher authority, and safeguard. All four connected a science-learning problem to an interaction and pedagogically interpretable evidence: misconceptions and gaps, explanations-in-progress, class-level readiness patterns, or investigation performance. However, only two made teacher control or evaluation explicit, and only two named a safeguard. The proposals therefore positioned AI less as an answer generator than as an elicitor, scaffold, and evidence-return mechanism, while leaving decision rights and protections unevenly specified. We argue that teacher professional learning should treat AI app ideation as epistemic specification work. A five-question design protocol--problem, learner interaction, evidence, teacher authority, and safeguard--can help teachers transform science-learning needs into accountable human-AI arrangements before building or adopting a tool.