Research
State Your Intention to Steer Your Attention: An AI Assistant for Intentional Digital Living
Overview Research area: Human-Computer Interaction (HCI), specifically digital self-control tools, attention management, and large language model (LLM)-driven proactive assistants. Published at CHI '2

- arXiv
- 2510.14513
- Published
- 2025-10-16
- Authors
- Juheon Choi, Juyong Lee, Jian Kim, Chanyoung Kim, Taywon Min, W. Bradley Knox, Min Kyung Lee, Kimin Lee
AI summary
Overview
Research area: Human-Computer Interaction (HCI), specifically digital self-control tools, attention management, and large language model (LLM)-driven proactive assistants. Published at CHI '26 (Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, April 13–17, 2026, Barcelona, Spain).
Technical level: Intermediate. The system design and evaluation are accessible, but the paper assumes familiarity with LLM prompting, classification metrics, and within-subjects field study methodology.
Scope: The paper introduces the Intent Assistant (INA), an LLM-based desktop assistant that elicits a user's intention, monitors on-screen activity to detect distraction, delivers gentle nudges, and refines its detection through user feedback—evaluated both on a purpose-built dataset (IntentionBench) and in a three-week field deployment with 22 participants.
What This Paper Is About
People using computers for study or work frequently drift away from what they intended to do, and most existing tools (blockers, time trackers, static reminders) rely on rules that cannot tell whether the same app is being used for a legitimate purpose or as a distraction. The authors build INA, an AI assistant that asks users to state their intention, clarifies it through a short Q&A, then continuously judges whether current on-screen activity matches that intention and intervenes with polite notifications when it does not. The goal is to test whether this kind of context-aware, feedback-refined assistant actually keeps people on task better than simply stating and periodically reviewing one's intention.
Key Contributions
- Formative study identifying design requirements. Interviews with 11 frequent computer users (3 undergraduates, 4 graduate students, 4 office workers, ages 20–30, average age 25.6) surfaced the limitations of static, rule-based productivity tools and informed four design goals: intention understanding, context-aware distraction detection, timely and gentle interventions, and feedback-driven refinement.
- The Intent Assistant (INA) system. An implemented AI assistant that elicits and clarifies a user's intention via LLM Q&A, monitors screenshots, application titles and URLs every two seconds, computes a distraction score, and issues dismissible nudges or praise.
- IntentionBench, a publicly released dataset. Built from 50 unique task instructions spanning 14 applications and 32 websites, yielding 350 mixed sessions of roughly 13 minutes each, totaling 138,803 data points and approximately 77 hours of activity, with realistic on-task/off-task transitions.
- A three-week in-the-wild field study. A within-subjects deployment with 22 analyzed participants comparing INA against a "simple reminder" application and a "logging only" application, combining quantitative measures with weekly surveys and interviews.
Main Findings
- Detection accuracy on IntentionBench: With full deployment configuration (clarification plus feedback), INA achieves accuracy 0.878, precision 0.959, recall 0.755, and F1 0.845.
- Both interactive features matter: Without clarification or feedback, accuracy is 0.805 and F1 0.739. Clarification alone yields accuracy 0.871 and F1 0.836; feedback alone yields accuracy 0.845 and F1 0.794. Clarification produces larger gains than feedback except on precision, but the best balance requires both.
- Real-world detection holds up: On a dataset curated from actual deployment records, INA attains an accuracy of 0.899, indicating it remains effective under noisy, ambiguous real usage.
- Reduced off-task behavior: Participants using INA showed a significantly lower LLM-estimated off-task ratio (0.104 vs. 0.166 with the simple reminder, p < .001).
- Higher intention alignment: Self-reported intention alignment ratings were 4.44 with INA vs. 4.23 with the simple reminder (p < .001), on a five-point scale.
- Greater focused immersion: INA scored 3.74 vs. 3.34 with the simple reminder (p = .045) and vs. 2.90 with logging only (p = .0003).
- Perceived benefits from qualitative data: Weekly surveys and interviews showed participants viewed INA as enhancing intention fulfillment, strengthening awareness of their digital habits, and providing supportive companionship.
- Acknowledged problems: Participants reported notification burden, imperfect detection accuracy, and challenges in long-term user adaptation.
Methodology in Plain English
The authors started by interviewing people who use computers heavily, to learn why existing blockers and reminders fail and what an assistant should do instead. They used those interviews to set four design goals.
They then built INA around an LLM. A session begins when the user types an intention, such as "study." The LLM asks two clarifying questions (the user can skip them or stop early) to turn a vague goal into something concrete. During the session, every two seconds the system sends the LLM the clarified intention, a screenshot of the current screen, and metadata about the focused application (title, and URL if it is a browser). The LLM returns a distraction score from 0 to 1, where 0 means perfectly aligned with the intention and 1 means fully misaligned; scores below 0.5 are treated as on-task. A notification is only sent if a change in state persists for 4 seconds, so brief glances do not trigger it. A switch to off-task triggers a polite, dismissible question inviting the user back to their intention; a switch back to on-task triggers praise. If the user stays off-task, the nudge repeats every 30 seconds.
Users can hover over any notification and mark it correct or incorrect, optionally adding a reason. When a notification is marked incorrect, the system records the screenshot, application or URL, and score, asks the LLM to explain why the detection was wrong, and appends a short refinement to the prompt for the rest of the session. Feedback is retained for a day to keep prompts from growing unbounded.
Because collecting real user data at scale is costly and because genuine on-task-to-off-task transitions are rare, the team constructed IntentionBench. Two researchers executed 50 distinct task instructions across 14 applications and 32 websites, capturing screenshots every second and performing a clarification Q&A beforehand to mimic the real pipeline. Each focused session was cut into segments at natural boundaries such as app switches. Mixed sessions were then synthesized by taking two focused sessions, concatenating them, and randomly reordering the segments; the intention from the first session labels its segments on-task and the second session's segments off-task. This yielded 350 mixed sessions averaging about 13 minutes each.
For the field study, 81 people completed a pre-survey, and participants were selected if they used a MacBook, were not employed at a corporation, and reported at least moderate digital distraction. Of 24 who began, 2 withdrew, leaving 22 analyzed (14 women, 8 men, aged 19–39; 9 aged 19–24, 8 aged 25–29, 5 aged 30–39; average self-reported computer use 5.6 hours per day, SD = 3.02). After a 30-minute online orientation, each participant used all three applications for seven days each in randomized order, with application names masked as color labels (Purple, Blue, Orange). Participants were asked to use their computer at least two hours per day with the assigned app running. Each week ended with a post-survey and a 10-20 minute semi-structured video interview.
The system itself was built in Python: a native macOS client with PyQt and a FastAPI backend on Google Cloud Platform, using Gemini 2.0-Flash at temperature 0.1. Screenshots were masked with Cloud DLP before upload, masked images stored in Cloud Storage, and metadata and event logs kept in Firestore, with original screen images used for real-time inference immediately discarded.
Why This Matters
Impact on research: The paper shifts digital self-control away from static rules and toward semantic alignment between a stated intention and observed on-screen context, and it supplies a reusable benchmark (IntentionBench) with realistic transition points for evaluating such systems. It also demonstrates a design pattern—clarification plus in-session feedback—that measurably improves LLM classification, with clarification contributing the larger gain and both features needed for the best result.
Real-world applications:
- Study and deep-work sessions, where students declare a goal and the assistant detects when browsing drifts into unrelated content.
- Knowledge work such as coding or document writing, where the same apps (browsers, video platforms, email) are legitimately used for task work, so blacklist tools fail.
- Wellbeing and attention-management software that nudges rather than blocks, preserving user autonomy.
- Evaluation infrastructure: IntentionBench and the released source code (https://intentassistant.github.io) can be used by other researchers to benchmark distraction-detection models.
Industry relevance: The findings are directly relevant to teams building productivity, focus, or wellbeing features into desktop operating systems, browser extensions, or LLM-based assistants, and to vendors weighing context-aware interventions against the notification burden they introduce. The measured effect sizes and the reported privacy architecture also matter for anyone deploying continuous screen monitoring in a commercial product.
Future Directions
- Reduce notification burden. Participants reported notification burden as a practical challenge, so determining an optimal intervention frequency and persistence remains open, despite the current design's choice of repetition every 30 seconds and a 4-second state-persistence gate.
- Improve detection accuracy further. Even the deployed configuration has recall of 0.755 on IntentionBench, so missed distractions and false alerts are unresolved; feedback is currently retained for only one day, and its long-term value is not established.
- Support long-term adaptation. Participants flagged long-term user adaptation as a challenge, raising the question of whether repeated feedback could build durable user-specific models rather than session-scoped prompt refinements.
- Generalize beyond the studied population. The field study analyzed 22 MacBook-using students and job seekers aged 19–39 who reported at least moderate distraction, so effectiveness for other populations, platforms, and work contexts is not reported.
Target Audience
HCI researchers and interaction designers working on digital self-control, attention management, or proactive AI; LLM practitioners interested in using semantic alignment for behavior classification with human-in-the-loop correction; and product teams building focus, productivity, or wellbeing features that must balance effective intervention against user autonomy and notification fatigue.
Authors’ abstract
When working on digital devices, people often face distractions that can lead to a decline in productivity and efficiency, as well as negative psychological and emotional impacts. To address this challenge, we introduce a novel Artificial Intelligence (AI) assistant that elicits a user's intention, assesses whether ongoing activities are in line with that intention, and provides gentle nudges when deviations occur. The system leverages a large language model to analyze screenshots, application titles, and URLs, issuing notifications when behavior diverges from the stated goal. Its detection accuracy is refined through initial clarification dialogues and continuous user feedback. In a three-week, within-subjects field deployment with 22 participants, we compared our assistant to both a rule-based intent reminder system and a passive baseline that only logged activity. Results indicate that our AI assistant effectively supports users in maintaining focus and aligning their digital behavior with their intentions. Our source code is publicly available at https://intentassistant.github.io