Research
From Knowledge Access to Source Learning: Developing Source-Specific Competence
Overview Research area: Natural language processing, specifically LLM agents that answer questions or act against an external body of knowledge (document collections, code repositories, API documentat

- arXiv
- 2610.02150
- Published
- 2026-10-01
- Authors
- Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang
AI summary
Overview
- Research area: Natural language processing, specifically LLM agents that answer questions or act against an external body of knowledge (document collections, code repositories, API documentation), plus retrieval-augmented generation and agent-memory systems.
- Technical level: Advanced. The paper is framed around formal problem statements, a persistent "source model," and mechanisms for grounded reconstruction; it assumes familiarity with retrieval, long-context modeling, and agent memory.
- Scope in one sentence: The paper defines "source learning," a way for an LLM agent to progressively build reusable, source-grounded competence about a persistent authoritative source, and evaluates the proposed SourceLearn system across five benchmarks and three LLM backends.
What This Paper Is About
Existing systems make external knowledge easier to retrieve or organize, and agent-memory systems preserve knowledge from past interactions, but repeatedly using the same source is still largely treated as repeated access rather than a chance to understand that source better. The authors define source learning: developing reusable source-specific competence over a persistent, authoritative source — how its knowledge is structured, interpreted, and applied, including relationships among concepts, conditions, procedures, and how information combines. They propose SourceLearn, which maintains a persistent, revisable source model and refines it through two complementary mechanisms, while always reconstructing persistent updates from the original source (the source remains the factual authority).
Key Contributions
- Source learning formulation and representation. The paper formulates source learning as the development of reusable source-specific competence and introduces the source model as an explicit, persistent, and revisable representation of that competence, complementing direct source access while preserving the source as the factual authority.
- Two complementary learning mechanisms. Self-Directed Source Learning proactively studies the source to deepen incomplete understanding and connect fragmented knowledge, while Task-Guided Source Learning uses deficiencies and recurring demands revealed through downstream tasks to refine that understanding further.
- A grounding invariant. All persistent updates follow a shared principle: learning signals decide what should be reconsidered, while persistent knowledge must be reconstructed from the authoritative source. Reconstruction is committed only when supported by source evidence, and observations are never copied directly into persistent memory.
- Empirical evidence across source types. SourceLearn is evaluated on document QA, code QA, tool use, and interactive environments, reportedly achieving the best result in 13 of 15 settings, with gains over Hybrid RAG of up to 22.6 points.
Main Findings
- Best-in-class across most settings: Across five benchmarks and three LLM backends, SourceLearn achieves the best result in 13 of 15 settings and the second-best in one more.
- Average improvement over Hybrid RAG: SourceLearn improves over Hybrid RAG by +14.3, +4.9, and +13.4 points on average with GPT-5.6-Luna, gpt-oss-120b, and DeepSeek-V4.1-Flash respectively.
- Largest single gain: The largest reported gain over Hybrid RAG is +22.6 points on APIBench with GPT-5.6-Luna, where SourceLearn reaches 70.4 against Hybrid RAG's 47.8.
- Other headline results with GPT-5.6-Luna: SourceLearn scores 77.6 (+12.7) on MultiDoc2Dial, 80.2 (+14.2) on NarrativeQA, 69.5 (+12.6) on SWE-QA, and 81.5 (+9.5) on AppWorld.
- Results with DeepSeek-V4.1-Flash: SourceLearn scores 81.3 (+18.0) on MultiDoc2Dial, 76.2 (+12.0) on NarrativeQA, 61.6 (+15.7) on SWE-QA, 82.2 (+15.1) on APIBench, and 89.9 (+6.0) on AppWorld.
- One setting below the baseline: With gpt-oss-120b on AppWorld, SourceLearn scores 27.4, which is 3.0 points below Hybrid RAG's 30.4.
- Counterparts are less consistent: RAPTOR, HippoRAG 2, and AWM improve more unevenly across benchmarks, which the authors read as evidence that progressively learning the source transfers more consistently than static source structuring or retaining task experience alone.
- Both learning stages matter: In the ablation with GPT-5.6-Luna, the full model M_T is the most accurate on all three QA benchmarks; omitting either Self-Directed or Task-Guided Source Learning lowers accuracy, and the initial source model M_0 is the lowest.
- Both task-guided pathways matter: Removing either failure-guided local refinement or cross-task representation learning lowers accuracy on every benchmark, and removing both (recovering M_S) lowers it further.
- Representation shifts away from look-up content: From M_0 through M_S to M_T, the representation shifts away from isolated look-up content toward structured rules and procedures.
- Conditions become explicit: Units with explicit applicability conditions rise from roughly 30% in M_0 to 65–82% in M_T across sources.
- Coverage of required source knowledge rises: The share of source claims required by test questions that are represented in the model rises from 23.2% in M_0 to 40.0% in M_T.
- Coverage tracks accuracy: Accuracy rises from 67.5% when none of the required claims are represented to 86.6% when all are represented.
- Not just compression: The authors argue these shifts show the source model reorganizes what is represented and becomes more useful for future tasks, rather than merely retaining or compressing more content.
Methodology in Plain English
The system treats one enduring knowledge source — a set of documents, a code repository, or an API library — as something the agent studies over time, rather than something it merely looks things up in.
Building the initial model. The agent reads the source in a source-aware, entity-centered way. Large sources are partitioned into coherent regions (or native units such as code modules, symbols, or document sections). For each region it builds representations of the major entities and the reusable structure spanning them, following a generic modeling instruction that favors major entities, important relations, procedures, governing conditions, distinctions, and representative details while avoiding redundant low-level content. The result, M_0, is deliberately provisional. It is not meant to reproduce the whole source; easily retrievable low-level details stay in the original source.
Self-Directed Source Learning. Each cycle follows an Inspect–Study–Consolidate process. The agent rereads an entity's source content alongside its current representation and records what the representation still fails to explain well — a partially captured mechanism, an implicit governing condition, an unintegrated reference. An adaptive planner then selects a small set of study actions at each round: Deepen (investigate an unresolved aspect of one entity) or Connect (jointly study two entities whose dependency or shared structure is insufficiently understood). Each action carries a focused temporary study question that guides renewed reading; these questions are learning instruments and are not stored. Accumulated observations are consolidated through grounded reconstruction, producing a "read-many, write-once" pattern.
Task-Guided Source Learning. Guidance tasks (a disjoint set from the test tasks) are used to expose two levels of deficiency. For each guidance task, the system returns to the source to identify what evidence supports the reference outcome and what source understanding the task requires. This is the only stage that directly observes the reference answer, and it uses it only to locate authoritative support and derive requirements — neither the answer nor its supporting evidence is stored as persistent knowledge. If a guidance task is answered incorrectly, the requirements absent from the current model are identified as learning targets and the implicated source region is reread and revised by grounded reconstruction, producing a locally refined model (M_local). Separately, each guidance task yields a "representation lesson" — a preference about how source knowledge should be represented, such as keeping distinctions explicit or not collapsing conditions — and recurring lessons are aggregated into a source-level representation policy that does not itself add source facts. The model is then recalibrated against the source under that policy, producing M_T.
Grounded reconstruction, shared by all stages. For an affected source region, the system inspects the source content against the current model to produce temporary observations, then reconstructs and applies an update. A reconstruction is committed only when the proposed content is supported by source evidence and previously supported meaning is preserved unless subsumed by a stronger representation.
Activation at task time. If the whole source model fits within a budget of B tokens, it is used whole; otherwise an activation routine selects task-relevant regions while preserving their internal structure. The solver then receives both the activated source understanding and task-specific evidence retrieved directly from the authoritative source, so the two contexts play complementary roles.
Evaluation setup. Five benchmarks span document, code, and API sources: MultiDoc2Dial, NarrativeQA, and SWE-QA for source-grounded QA; APIBench for tool use; and AppWorld for interactive tasks. Backends are GPT-5.6-Luna, gpt-oss-120b, and DeepSeek-V4.1-Flash. Guidance and test tasks are disjoint and shared across methods, using three fixed 30%/70% splits except AppWorld, which uses its official training and test-normal splits. QA answers are judged by GPT-5.6-Luna, while APIBench and AppWorld use their official evaluators. Counterparts are Hybrid RAG, RAPTOR, HippoRAG 2, and AWM, with all methods sharing the embedding model (text-embedding-3-large) and, on AppWorld, the same agent scaffold.
Why This Matters
- Research impact: The paper reframes repeated interaction with the same external source as a learning problem about the source itself, distinguishing it from source-access methods that optimize query-time use and from agent-memory methods that preserve interaction-derived knowledge. It argues that these optimize different objects, and that persistent updates should be reconstructed from an authoritative source rather than stored from task experience.
- Real-world applications:
- Long-lived assistants over organizational document collections, such as government-domain documentation spanning DMV, SSA, StudentAid, and VA as in MultiDoc2Dial.
- Developer agents working repeatedly against large code repositories, tested here on streamlink, sphinx, xarray, pytest, flask, and conan in SWE-QA.
- Tool-use agents that must reason about dependencies and prerequisites across API calls, as in APIBench's TorchHub, HuggingFace, and TensorFlow Hub documentation.
- Interactive agents operating against simulated app APIs, as in AppWorld's nine apps and 457 API specifications.
- Industry relevance: Deployments where many tasks depend on one stable knowledge asset — internal documentation portals, codebases, and API catalogs — could benefit if the agent's understanding of that asset keeps improving rather than being reconstructed task by task. The approach also keeps the original source as the factual authority, which matters where traceability to authoritative evidence is required. AWM-style workflow induction alone improved unevenly across benchmarks, and in one reported setting (gpt-oss-120b on AppWorld) SourceLearn itself scored 3.0 points below Hybrid RAG, so gains are not uniform in every configuration.
Future Directions
- Extending beyond stable sources: The authors state that the current scope is limited to persistent, authoritative, and relatively stable sources with repeated same-source tasks; extending source learning to evolving, noisy, or conflicting sources remains an important direction.
- Understanding guidance-task budget: The reported learning uses a 30% guidance split (or AppWorld's 90 official training tasks); how much guidance is needed for source learning to pay off is not reported in the provided content.
- Why some configurations lag: The one setting where SourceLearn fell below Hybrid RAG (gpt-oss-120b on AppWorld) is not explained in the provided content, leaving open which source and backend properties govern when source learning helps.
- Interaction with richer source representations: The paper compares against static representations such as RAPTOR and HippoRAG 2 and against experiential memory such as AWM, but does not report combining source learning with those mechanisms.
Target Audience
Researchers and advanced practitioners working on LLM agents, retrieval-augmented generation, long-context knowledge use, and agent memory; engineers building assistants that repeatedly operate over codebases, documentation collections, or API catalogs; and readers interested in how agents can accumulate reusable, source-grounded competence rather than re-deriving the same understanding on every task.
Authors’ abstract
Large language model (LLM) agents increasingly rely on persistent external sources to solve sequences of knowledge-intensive tasks. Existing methods improve how source content is accessed and organized, while agent-memory systems preserve reusable knowledge from prior interactions, but repeated use of the same source is still largely treated as repeated access rather than an opportunity to progressively improve understanding of that source. We study source learning: developing reusable source-specific competence over a persistent authoritative source. We represent this competence with a persistent source model that captures reusable understanding of the source, including how its knowledge is structured, interpreted, and applied. To construct and progressively refine such models, we propose SourceLearn, which combines two complementary learning mechanisms. Self-Directed Source Learning identifies what remains incompletely understood and adaptively revisits the source, while Task-Guided Source Learning uses downstream experience to reveal local representational gaps and recurring needs in how source knowledge should be organized. In both cases, learning signals determine what should be reconsidered, while persistent updates are reconstructed from the authoritative source. Across five benchmarks and three LLM backends, SourceLearn achieves the best performance in 13 of 15 settings, with gains of up to 22.6 points over Hybrid RAG and substantial overall improvements over static source representations and experience-based memory baselines.