Skip to main content Skip to secondary navigation

Education Data Science Conference Agenda

Main content start

Agenda: May 27 

8:30-9:15amRegistration and Light Breakfast 
9:15-9:30am 

Welcome Remarks 

9:30-10:15am

Keynote Talk: Practice-based Educational Data Science: A Proposal for Our Emerging Field

Joshua Rosenberg

In this talk, Joshua Rosenberg examines whether educational data science is primarily a methodological program that takes education as a use case, or work situated in educational contexts and accountable to educational outcomes. Arguing that these paths lead to different kinds of work,  he will highlight key features of practice-based approaches, consider implications for teacher education and tools, and reflect on how the field can maintain rigor while better serving educators and learners in an era of AI.

10:15-11:15am

Panel I: Annotation 

This session brings together three perspectives on data annotation—from LLMs as high-throughput labelers to human–AI codebook co-design and culturally situated (“glocal”) coding—to compare what “annotation” means across approaches. We will examine how codebooks are developed and validated, and we will interrogate reliability by asking when disagreement reflects noise versus real ambiguity, shifting norms, or context-specific meaning—and what that implies for responsible, transparent use of LLMs in annotation pipelines.

  • Sarah Williams-Habibi: Coding with Claude: Recommendations for Exploring AI-Assisted Qualitative Analysis in Mixed-Methods Education Research
  • Danielle Thomas: Consensus is Not Correctness: Rethinking Reliability in Educational AI
  • Zhuqian Zhou: Optimizing LLM Annotation of Classroom Discourse through Multi-Agent Orchestration
  • Moderator: Dora Demszky
11:15-11:30amBreak
11:30am-12:05pm 

Lightning Talks: AI Use in Education

  • Ran Liu: AI for All? Uncovering Sociodemographic Gradients in Student-AI Interactions using a Mixed-Methods Framework
  • Sina Rismanchian: Less Time For Learning: Evidence of Students Resorting to GenAI in an Adaptive Learning and Assessment System
  • Lief Esbenshade: Midyear Evidence on Teacher's Use of Generative AI in K-12 Schools
  • Renzhe Yu: How Have Instructors Adapted to Generative AI? Evidence from 90,000 Assignments in Postsecondary Courses
  • Jinsook Lee: The Digital Divide with Generative AI: Examining LLM Adoption and Lexical Convergence Across Socioeconomic Groups in College Admissions Essays
  • Mei Tan: Feedback Footprint: Modeling the Language of Written Feedback Across Teachers and LLMs
12:05pm-12:30pm

Poster Session I

  • Maxime Lelièvre: Talking teachers' languages testing AI multilingual pedagogy ability
  • Kirk Vanacore: Establishing Baselines for Foundation Models in Educational Discourse Analytics
  • Zhen Xu: Enhancing LLM-Based Data Annotation with Error Decomposition
  • Echo Zexuan Pan: Evaluating Large Language Models as Scoring Tools for Children's Responses to Open-Ended Vocabulary Questions
  • Shelby Smith: Urban AI Unlocked: How District Procurement and Spending Shape Equitable AI Use in K-12 Schools
  • Puja Maharjan: Evaluating Automated Teaching Quality Assessment: ASR Performance, Diarization Evaluation, and Downstream Prediction
  • Ran Liu: AI for All? Uncovering Sociodemographic Gradients in Student-AI Interactions using a Mixed-Methods Framework
  • Sina Rismanchian: Less Time For Learning: Evidence of Students Resorting to GenAI in an Adaptive Learning and Assessment System
  • Lief Esbenshade: Midyear Evidence on Teacher's Use of Generative AI in K-12 Schools
  • Renzhe Yu: How Have Instructors Adapted to Generative AI? Evidence from 90,000 Assignments in Postsecondary Courses
  • Jinsook Lee: The Digital Divide with Generative AI: Examining LLM Adoption and Lexical Convergence Across Socioeconomic Groups in College Admissions Essays
  • Mei Tan: Feedback Footprint: Modeling the Language of Written Feedback Across Teachers and LLMs
12:30-1:30pmLunch 
1:30-2:00pm

Lightning Talks: Describing Policy, Experiences, and Curriculum

  • Jose Aguilar: From Public Comment to Policy Signals: Validating Interpretable NLP for Curriculum Reform
  • Michael Chrzan: The Language of Closure: Examining Racial Differences in How A Community Discusses School Closure Metrics
  • Yan Jiang: Centering Lived Experiences in Education Data Science: A Five-Year Analysis of Early Care and Education Provider's Experiences with Challenges
  • Jing Liu: The Production and Consequences of Racial Disparities in School Discipline
  • Jake Nicoll: (How) do we teach emotions?
2:00pm-2:30pm

Poster Session II

  • Chelsea Brown: Talk Trees: Supporting Student Agency in Classroom Discourse
  • Immanuel Williams: Integrating Workflow-Centered and Library-Centered Instruction in Data Science Education
  • Wei Wang: Bridging Workforce Data and Decision-Making with a RAG-Enabled Community College Data Explorer
  • Lily Roth: An Exploration of How the United States Conceptualize Data Science in K-12 Education
  • Irakli Matcharashvili: Early Childhood Education as a Long-Term Investment. Heterogeneous Associations Across Global Income Contexts
  • Tsubasa Matsuoka: Mapping Long-Run Shifts in Japan's English Language Education Research: Structural Topic Modeling of ARELE Abstracts (1990-2024)
  • Tom Nachtigal: Automating Education Reform Identification with Large Language Models: Validity and Scale in World Education Reform Data
  • Annaliese Paulson: Classifying Courses at Scale: Using Large Language Models to Standardize Postsecondary Transcripts and Identify Curricular Deserts
  • Jose Aguilar: From Public Comment to Policy Signals: Validating Interpretable NLP for Curriculum Reform
  • Michael Chrzan: The Language of Closure: Examining Racial Differences in How A Community Discusses School Closure Metrics
  • Yan Jiang: Centering Lived Experiences in Education Data Science: A Five-Year Analysis of Early Care and Education Provider's Experiences with Challenges
  • Jing Liu: The Production and Consequences of Racial Disparities in School Discipline
  • Jake Nicoll: (How) do we teach emotions?
2:30-2:45pmBreak
2:45-3:45pm

Panel II: Applied Lessons 

This session highlights practical lessons for designing, implementing, and evaluating data-driven interventions in education. From cross-sector data partnerships that enable early risk prediction, to scalable causal frameworks within adaptive tutoring systems, to randomized evaluations of early childhood language interventions, these papers demonstrate how rigorous methods can be embedded in real-world contexts to generate actionable insights for educators and policymakers.

  • Amy Burkhardt: Evaluating the Theory of Action for an NLP-Based Writing Feedback Tool: Educator Reactions and Student Writing Behaviors
  • Jessica Rood: Beliefs, Conversations, and Language Development in Early Childhood: Evidence and Mechanisms from an Experimental Study
  • Kirk Vanacore: TutorImpact: A Causal Framework for Estimating Skill-Level Effects of On-Demand Tutoring in Adaptive Learning Systems
  • Moderator: Brian Kim, Founder & Chief Data Scientist, Open Augments
3:45pm-4:30pm

Small Group Discussions 

  • Augmenting Human Analytic Power with Machine Learning (Facilitated by Christina Krist)
  • Rethinking Teaching in the Age of AI (Facilitated by Emma Brunskill)
  • Education And the Changing Labor Market in The Age of AI (Facilitated by Mitchell Stevens)
  • Recent Advances in Psychometric and Measurement Research (Facilitated by Ben Domingue)
  • Building Research–Practice Partnerships in Education Data Science (Facilitated by Laura Wentworth)
  • Designing Experiences that Support Learning (Facilitated by Karin Forssell)
4:30-4:35pmClosing Remarks 
4:35-5:30pm

Networking Reception & EDS MS Posters

  • Ari An: Differentiated SDG4 Policy Agendas: A Global Analysis of Education Reform Portfolios, Observed Deficits, and Temporal Differentiation, 2000–2021
  • Vryan Feliciano: "Who’s Coding What?" An Analysis of YouTube Computer Science Education Videos
  • Kazunori Fukuhara: From Persona to Parameters: Controlling LLM-Simulated Students with IRT and CDM
  • Xinman Liu: Beyond Automation and Augmentation: A Mixed-Methods Study of Teacher-AI Collaboration in Feedback Provision
  • Andrea Mock: Behavioral and Temporal Effects of Digital Wellbeing Interventions Among Adolescents
  • Savira Nadela: Precision vs. Prediction: The Relationship between Parametric Standard Errors and Prediction Quality
  • Mitsutoshi Nozaki: The Impact of California's Accountability Incentives and Supports on Chronic Absenteeism
  • Mayank Sharma: CONVOLEARN: A Dataset for Fine-Tuning Dialogic AI Tutors
  • Yimei Shen: Family and School Learning Ecologies: Pathways to Adolescents’ Sexual Self-Efficacy in the 1990s United States
  • Teah Shi: Evaluating Open-Source LLM Adaptation for Belonging-Centered Instruction: Detecting and Rephrasing Teacher Utterances in Elementary Math Classrooms
  • Ziqi Shu: Measuring Optimal Challenge: Trajectory-based Difficulty Alignment in Open-Ended Language Tutoring
  • Jason Zhang: Themes and Sentiment in Alumni Reflections: A Multilevel Analysis of Associations with Civic, Psychological, and Career Outcomes

Agenda: May 28

8:30-9:15amRegistration and Light Breakfast 
9:15-9:30am Welcome to Day 2
9:30-10:15am

Keynote Talk: Where Does Judgment Happen? Educational Data Science in the Age of Generative AI

Danielle McNamara

In the age of generative AI, educational data science faces a new challenge: large language models increasingly generate explanations, recommendations, and interpretations that function as judgments rather than simply supporting analysis. Across themes of annotation, multimodal data, equity, validation, intervention, and AI-assisted analysis, McNamara argues that AI compresses intermediate reasoning steps, making it easier for plausible outputs to shape educational decisions before their basis is fully established. Alongside prediction, measurement, and design, she proposes stewardship as a fourth paradigm for educational data science, focused on governing how analytic outputs are interpreted, validated, and used in practice. The talk outlines a framework for stewardship grounded in epistemic discipline, provenance, accountability, institutional learning, and learner agency.

10:15-11:15am

Panel III: Multimodal 

This session features studies using multiple sources of data (audio traces, video traces, writing traces) to better understand and/or predict educational phenomena. The studies focus on listening skills, productive struggle, and speaker attribution, respectively and all argue that education research benefits from picking up traces of learning given that learning is a very complex process that is expressed in multiple ways. The deeper question in this session is how multimodal learning analytics can provide better insight into issues of learning and teaching

  • Michael Chrzan: Multimodal Speaker Identification in Classroom Environments
  • Xiaomeng Huang: Towards Automated Measurement of Active Listening Skills in Collaborative Problem Solving Using Multimodal Learning Analytics and Large Language Models
  • Wanjing Anya Ma: The Role of Whiteboard Data in Assessing Tutoring Quality
  • Moderator: Peter Alexander Youngs, Professor and Chair, UVA Department of Curriculum, Instruction & Special Education
11:15-11:30amBreak
11:30am-12:00pm 

Lightning Talks: Equity and Bias

  • Jinwon Kim: Toward More Equitable Learning Environments: Insights from Digital Trace Data on Inclusive Course Design Features
  • Katharine Sadowski: From Fragmented Data to Actionable Insight: Predicting Asthma Risk Through a School's Health Data Partnership
  • Mayank Sharma: Beyond Blind Spots: Mitigating Intersectional Bias in Student Dropout Prediction
  • Dong Chen: Evaluating Racial Bias in Large Language Models for College Admissions Scoring
  • Zhen Xu: Intersectional inequalities in asynchronous academic communication
12:00pm-12:30pm

Poster Session III

  • Jeffrey Bush Tayne: A Co-Design Intervention with AI analytics to Support the Academically Productive Talk of Title 1 School Math Tutors and their Students: A Comparative Interrupted Time Series Analysis
  • Roland Molontay: Predicting Dropout with Explainable AI Using ProgressivelyAvailable Student Data
  • Wei Gao: Disrupted Learning: The Effects of Within-Year School Transfers on Academic Progress
  • Chengyuan Yao: Attending to Distribution Shifts: Decomposing Transfer Performance of Educational Predictive Models
  • Xin Wei: Triangulating Discovery: How Multi-Source Retrieval Augmentation Improves Variable Discovery in Educational Data Archives
  • Marcia Yang: Network Community Detection Methods for Fuzzy Matching Across Multiple Datasets
  • Margeaux Randolph: Designing with Generative AI in Early Learning Contexts: Evidence on Representation, Validity, and Ethical Constraints
  • Jinwon Kim: Toward More Equitable Learning Environments: Insights from Digital Trace Data on Inclusive Course Design Features
  • Katharine Sadowski: From Fragmented Data to Actionable Insight: Predicting Asthma Risk Through a School's Health Data Partnership
  • Mayank Sharma: Beyond Blind Spots: Mitigating Intersectional Bias in Student Dropout Prediction
  • Dong Chen: Evaluating Racial Bias in Large Language Models for College Admissions Scoring
  • Zhen Xu: Intersectional inequalities in asynchronous academic communication
  • Bo Pei: Support Statistical Learning through Platform Angles: A Mixed-Methods Study on Scaling Learning Insights in Higher Education Statistics
12:30-1:30pmLunch 
1:30-2:30pm

Panel IV: Validation

This session positions validity as more than a question of model performance, instead examining how educational constructs are defined, operationalized, and represented through data science methods. Across statistical modeling, computational analysis of curricular content, and AI-based interpretation of teacher narratives, the papers highlight how methodological choices shape not only results, but the very phenomena that become visible and measurable in educational research.

  • Shayan Doroudi: Students (Probably) Don't All Learn at the Same Rate: A Tale of Model Misspecification
  • Emileigh Harrison: Separation of Church and State Curricula? Examining Public and Religious Private School Textbooks
  • Mingyan Ma: Using Large Language Models to Analyze Beginning Teacher Challenge Narratives
  • Moderator: Katharine Sadowski
2:30-2:45pmBreak
2:45pm-3:35pm

Workshop: Accelerating Rigorous Education Research with AI Agents: An Introduction to the Data Analyst Augmentation Framework (DAAF)

AI agents can now autonomously plan, write, review, and execute analytic code, raising urgent questions about their role in research given known risks like hallucinations and inaccuracies. This session introduces DAAF, an open-source framework for Claude Code, to help researchers leverage these tools for data analysis while maintaining transparency, reproducibility, and rigor.

Brian Kim, Founder & Chief Data Scientist, Open Augments

3:35-4:35pm

Closing Panel: Industry Perspectives

What skills and mindsets does industry need from education data scientists and analysts today? Join for a candid closing conversation on staying up-to-date and preparing for where the field is headed.

4:35-4:45pmClosing Remarks 
4:45-6:00pmNetworking Reception