Research
Exploring Student Expectations and Confidence in Learning Analytics
Overview Research area: Learning Analytics (LA) adoption in higher education, specifically student attitudes toward data processing, data protection, and LA features — sitting at the intersection of a

- arXiv
- 2601.05082
- Published
- 2026-01-08
- Authors
- Hayk Asatryan, Basile Tousside, Janis Mohr, Malte Neugebauer, Hildo Bijl, Paul Spiegelberg, Claudia Frohn-Schauf, Jörg Frochte
AI summary
Overview
- Research area: Learning Analytics (LA) adoption in higher education, specifically student attitudes toward data processing, data protection, and LA features — sitting at the intersection of applied machine learning (clustering, decision trees) and educational technology.
- Technical level: Intermediate. The machine learning used (K-means, elbow and silhouette methods, CART decision trees) is standard, but readers benefit from familiarity with clustering and survey-based research.
- Scope in one sentence: An anonymous survey of 553 students across eight fields of study at a German university of applied sciences, using an adapted 12-item Student Expectations of Learning Analytics Questionnaire (SELAQ), is analyzed with clustering and decision-tree methods to identify four distinct student attitude profiles.
What This Paper Is About
Learning Analytics systems collect and analyze student data to understand and improve learning, but doing so raises data protection and privacy concerns, and such systems fail if students reject them. This paper asks how students across different degree courses actually feel about the collection, protection, and use of their educational data, and whether those feelings are uniform or split into recognizable groups. The goal is to give universities an empirically grounded picture of student acceptance so that LA implementations can be designed and communicated more effectively.
Key Contributions
- A comprehensive survey of university students about their attitudes and expectations regarding Learning Analytics systems, conducted in class during Fall 2022 and Summer 2023 to ensure a random sample from participating faculties.
- Discipline-level analysis of survey responses, distinguishing how student groups classified by academic discipline (Computer Science, Architecture, Civil Engineering, Electro-Mechanical Engineering, Sustainability, Surveying, Business Studies, and Other Engineering Courses) respond differently.
- Automatic clustering of the student population using K-means on the responses to 24 questions (12 desires and 12 expectations) from 553 students, yielding four clusters named Enthusiasts, Realists, Cautious and Indifferents.
- A decision-tree based explainability layer (CART) that models the clusters using Data Protection and Learning Analytics scores as features, plus a distribution analysis of clusters across fields of study to lay the foundation for successful and broadly accepted LA implementations.
Main Findings
- Four student clusters were identified. Using the elbow and silhouette methods, the optimal number of clusters was estimated as k = 4. The clusters are named Cluster A (Enthusiasts), Cluster B (Realists), Cluster C (Cautious) and Cluster D (Indifferents).
- Overall averages across all clusters: Data Protection desire 5.91, Data Protection expectation 5.39, LA desire 5.33, LA expectation 4.51 — expectations are always lower than desires.
- Enthusiasts (Cluster A) show LA desire 6.01 and LA expectation 5.64, with DP desire 6.34 and DP expectation 6.14 — high desire plus high confidence that the university will deliver.
- Realists (Cluster B) match Enthusiasts on desire (DP desire 6.34, LA desire 5.76) but drop sharply on expectation (DP expectation 4.95, LA expectation 3.73) — they want these things but doubt they will happen.
- Cautious (Cluster C) show the lowest LA desire and expectation (3.68 and 3.75) while maintaining robust DP desire and expectation (6.19 and 5.64) — skeptical about LA's effectiveness but firm about data protection.
- Indifferents (Cluster D) report relatively similar, medium-level ratings across the board: DP desire 3.55, DP expectation 4.02, LA desire 4.24, LA expectation 3.91 — a general lack of interest or involvement.
- Question 11 stands out as an outlier. The expectation that teaching staff have an obligation to act (support students at risk of failing or underperforming, or improving their learning) is significantly lower than for all other questions, suggesting students are not very convinced that their teachers provide support and optimize learning success.
- Lecturer-related items score lower. Both desire and expectation drop noticeably in the LA Lecturer context compared with the general LA items, with the effect stronger for expectations.
- Discipline differences are visible. Sustainability students give significantly higher scores on most questions, indicating generally high desires and expectations. Many Civil Engineering students rate below the average.
- Computer Science students showed stronger agreement with statements 2d (secure handling of educational data), 3e (consent before third-party analysis) and 9d (desire for a complete learning profile across modules) — possibly linked to their IT background.
- The decision tree used two effective split thresholds: grade 4.6 and grade 4.8, partitioning the DP and LA features into branches.
- Cluster distribution across fields: roughly 20% of Engineering and Computer Science students fall into Cautious and a notable 40% into Enthusiasts. Computer Science has fewer Indifferents and slightly more Realists. Architecture leans more toward Realist. A noteworthy number of Business Administration students fall into the Indifferent category. Sustainability students show a higher percentage in the Enthusiasts category.
Methodology in Plain English
The researchers ran an anonymous, in-class survey at one university in Fall 2022 and Summer 2023. The instrument was an adapted, translated version of the 12-item Student Expectations of Learning Analytics Questionnaire (SELAQ), extended to capture degree program information. Each of the 12 questions was answered twice: once for desire ("Ideally, I would like that to happen") and once for expectation ("In reality, I would expect that to happen"), each rated from 1 (strongly disagree) to 7 (strongly agree). That yields 24 numeric responses per student.
The 12 questions map to three groups: Data Protection (questions 1–3, 5, 6), LA General Functionality (questions 4, 7–9, 12), and LA Lecturer-related Features (questions 10, 11).
Missing values were imputed with the mean for numerical features and the most common value for categorical features. The researchers then averaged feature scores by group and by question, and checked which fields of study produced the lowest and highest averages.
To see whether students were homogeneous or formed groups, they applied K-means clustering to the 24 responses of 553 students. The elbow and silhouette methods both indicated k = 4, so K-means was rerun with that value. For each cluster they computed average scores by question group.
Finally, to make the clusters interpretable, they trained a CART decision tree using DP and LA scores as features to predict cluster membership. This produced a simple, transparent rule structure with split thresholds at 4.6 and 4.8. They also modeled the distribution of the four clusters across fields of study.
Why This Matters
Impact on research: Prior SELAQ studies had not segregated students by personal or university traits. This paper argues that recognizing discipline-specific nuances offers a more holistic understanding of student LA attitudes, and it claims to be the first to apply this clustering-and-decision-tree combination in this context. It also builds on findings from Europe (Kollom et al., 2021), Latin America (Garcia et al., 2021), Brazilian students (Pontual Falcão et al., 2022), and the 417 European students surveyed by Wollny et al. in 2023, adding a discipline-aware perspective.
Real-world applications:
- Tailored rollout strategies: Universities can adjust LA communication and deployment based on the dominant cluster in a given faculty — reassuring Realists with strict DP standards and transparent communication, engaging Cautious students with workshops on LA benefits and concrete outcomes, and combining both approaches for Indifferents.
- Protecting the Enthusiasts: For students naturally inclined toward LA, the priority is guaranteeing robustness of LA systems and actually delivering on the promised features so these students are not lost.
- Lecturer training: The weak scores on lecturer-related items point to a need for specialized LA training for teaching staff and open discussion about LA's utility and the lecturer's role.
- Data protection messaging: Because DP desire and expectation remain high even in the Cautious cluster, institutions should foreground rigorous data protection as part of any LA pitch.
Industry relevance: Vendors building learning management systems, dashboards, and analytics products for higher education can use these profiles to design onboarding, consent flows, and feature emphasis that match how different student populations actually reason about data and analytics. The finding that expectations consistently trail desires suggests product and institutional messaging that only promises benefits without demonstrating delivery risks losing credibility.
Future Directions
- How can instructors effectively communicate with and address the concerns of cautious students, ensuring their trust in the use of LA?
- What strategies can enhance engagement of indifferent students and highlight the advantages LA can offer?
- Extending beyond technical subjects: Because participants were mainly from technical fields, future studies could test whether expectations and concerns of students from Humanities or Social Sciences differ significantly from these findings.
- Discipline-tailored institutional strategy: The paper suggests cluster proportions may vary by academic discipline, implying universities with pronounced specializations should recognize and adjust to their dominant student cluster — a hypothesis that would benefit from testing at other institutions.
Target Audience
This paper is most useful to Learning Analytics researchers and educational data mining practitioners who work with student-facing survey instruments such as SELAQ; university administrators, deans, and data protection officers planning or defending LA deployments; instructional designers and teaching staff who will be expected to act on analytics output; and edtech product teams building dashboards, consent mechanisms, and analytics features for higher education. Readers interested in interpretable machine learning applied to survey data will also find the clustering-plus-decision-tree combination instructive.
Authors’ abstract
Learning Analytics (LA) is nowadays ubiquitous in many educational systems, providing the ability to collect and analyze student data in order to understand and optimize learning and the environments in which it occurs. On the other hand, the collection of data requires to comply with the growing demand regarding privacy legislation. In this paper, we use the Student Expectation of Learning Analytics Questionnaire (SELAQ) to analyze the expectations and confidence of students from different faculties regarding the processing of their data for Learning Analytics purposes. This allows us to identify four clusters of students through clustering algorithms: Enthusiasts, Realists, Cautious and Indifferents. This structured analysis provides valuable insights into the acceptance and criticism of Learning Analytics among students.