Shopping for tiramisu ingredients
A1_JAKE · DAY1 · 17:32:00–17:33:00Context resolution: raw captions and elliptical speech omit what “it” refers to, while EgoCITE constructs a self-contained index for ordering the cake ingredients online.
Long-horizon egocentric memory requires agents to recover relevant experiences from days of continuous first-person video and audio. Existing systems typically store isolated video captions or speech transcripts and retrieve them using raw user queries. We find that this pipeline loses the situational context needed for reliable memory search: stored fragments are not self-contained, complementary descriptions of the same experience remain disconnected, and queries often express temporal intent only implicitly. We introduce EgoCITE, a situation-aware framework that contextualizes both memory storage and retrieval for long-horizon egocentric question answering. EgoScheme uses local multimodal context to transform fragmentary captions and transcripts into self-contained atomic memories. EgoIndex organizes complementary action, activity, utterance, and conversation representations into multi-view, multi-granularity indices that capture each situation from different perspectives. EgoRetrv interprets the situation implied by a question through question-conditioned temporal relevance scoring, then iteratively retrieves and curates matching evidence. We evaluate EgoCITE on EgoLifeQA, EgoMem, and EgoR1-Bench using answer accuracy and target-event retrieval alignment. Across the three benchmarks, EgoCITE improves average answer accuracy over the strongest agentic memory baselines by 4.4–14.2% while achieving 36× lower normalized cost than long-context LLM agents.
Context resolution: raw captions and elliptical speech omit what “it” refers to, while EgoCITE constructs a self-contained index for ordering the cake ingredients online.
Multi-view indexing: EgoCITE separates physical handling of cooking tools from spoken coordination while retaining the broader cooking-setup activity.
We evaluate EgoCITE on EgoLifeQA, EgoMem, and EgoR1-Bench in terms of answer accuracy and target-event retrieval alignment (retrieval hit rate).
| Method | Input tokens | Accuracy (%) | Retrieval Hit Rate (%) | ||||
|---|---|---|---|---|---|---|---|
| EgoLifeQA | EgoMem | EgoR1-Bench | EgoLifeQA | EgoMem | EgoR1-Bench | ||
| Long-context LLM agents | |||||||
| GPT-5.4 agent | 783k | 54.9 | 75.1 | 71.0 | — | — | — |
| Gemini-3.1-Pro agent | 804k | 60.2 | 78.3 | 72.7 | — | — | — |
| Agentic memory baselines | |||||||
| EgoRAG | 16k | 34.3 | 59.5 | 49.5 | 4.3 | 41.1 | 7.6 |
| A-MEM | 6k | 42.9 | 58.8 | 52.0 | 15.4 | 41.4 | 24.7 |
| VideoRAG | 21k | 46.5 | 68.6 | 64.7 | 2.0 | 17.4 | 6.0 |
| WorldMM-Qwen | 87k | 46.4 | 74.8 | 66.3 | 31.0 | 72.6 | 38.3 |
| WorldMM-GPT | 134k | 49.6 | 68.1 | 64.0 | 34.5 | 66.9 | 40.0 |
| EgoCITE-Qwen (ours) | 34k | 61.5 | 79.9 | 73.0 | 43.4 | 86.8 | 59.0 |
| EgoCITE-GPT (ours) | 32k | 63.8 | 79.2 | 75.7 | 49.6 | 89.6 | 62.7 |
Multiple-choice answer accuracy is evaluated across five EgoLifeQA question categories and four EgoMem question categories.
Result: EgoCITE-GPT achieves the highest accuracy in all nine categories, outperforming the strongest displayed agentic memory baseline in each category by 5.8–20.5 percentage points on EgoLifeQA and 3.4–15.3 percentage points on EgoMem.
Insight: These consistent gains show that EgoCITE’s structured memory representation and time-aware retrieval improve accuracy across diverse egocentric question types.
Cumulative multiple-choice accuracy on EgoLifeQA is measured as the available memory horizon expands from DAY1 through DAY7.
Result: EgoCITE-GPT leads every baseline from DAY3 onward and reaches 62.8% at DAY7, 2.5 percentage points above the next-best Gemini-3.1-Pro agent at 60.3%.
Insight: EgoCITE’s structured multi-view memory and time-aware retrieval scale more effectively than both long-context reasoning and existing retrieval-based approaches.
Move across the plot or use the day slider to compare all visible methods at the same memory horizon.
We evaluate the EgoLifeQA accuracy–efficiency trade-off using average retrieval latency per question and normalized cost, the GPT-5.4-price-weighted sum of input and output token costs, where lower latency and cost are better.
Result: Five-round EgoCITE-GPT reaches 63.8% accuracy at 32.0 s and 44.1k normalized cost, outperforming WorldMM-GPT by 14.2 percentage points with 14.8 s lower latency and GPT-5.4 by 8.9 points with 36× lower cost at comparable latency.
Insight: EgoCITE lies on the Pareto frontier because its concise, coreference-resolved atomic memory indices reduce input context per retrieved index by 5.7× relative to WorldMM captions, balancing accuracy, latency, and cost.
Higher accuracy and lower latency are better; bubble area encodes normalized cost. Select or hover a point for exact values.
Accuracy is evaluated on 2,223 EgoLifeQA questions containing explicit temporal cues, which constitute 76% of the benchmark.
Result: EgoCITE-GPT reaches 62.6% accuracy, outperforming Gemini-3.1-Pro by 3.9 percentage points and WorldMM-GPT by 15.5 percentage points.
Insight: EgoCITE captures temporal intent more effectively than both long-context LLM agents and existing agentic memory baselines.
| Method | Accuracy (%) Higher is better. |
|---|---|
| Long-context LLM agents | |
| GPT-5.4 agent | 51.6 |
| Gemini-3.1-Pro agent | 58.7 |
| Agentic memory baselines | |
| EgoRAG | 34.8 |
| A-MEM | 42.2 |
| VideoRAG | 43.4 |
| WorldMM-Qwen | 43.8 |
| WorldMM-GPT | 47.1 |
| EgoCITE | |
| EgoCITE-Qwen (ours) | 60.9 |
| EgoCITE-GPT (ours) | 62.6 |
Accuracy is evaluated on habitual questions in EgoLifeQA and multiple-evidence questions in EgoMem, both of which require retrieving information across multiple timestamps rather than from a single moment.
Result: EgoCITE-GPT reaches 64.3% accuracy on habitual questions and 80.3% on multiple-evidence questions, outperforming agentic memory baselines by 5.8–29.9 and 3.0–22.7 percentage points, respectively.
Insight: EgoRetrv aggregates and retrieves information across long temporal horizons while requiring more than 24× fewer input tokens than long-context LLM agents.
| Method | Accuracy (%) Higher is better. | |
|---|---|---|
| HabitualEgoLifeQA | Multiple evidenceEgoMem | |
| Long-context LLM agents | ||
| GPT-5.4 agent | 63.9 | 83.8 |
| Gemini-3.1-Pro agent | 59.8 | 81.2 |
| Agentic memory baselines | ||
| EgoRAG | 34.4 | 57.6 |
| A-MEM | 44.3 | 61.6 |
| VideoRAG | 58.5 | 70.7 |
| WorldMM-Qwen | 49.5 | 77.3 |
| WorldMM-GPT | 53.6 | 64.6 |
| EgoCITE | ||
| EgoCITE-Qwen (ours) | 58.8 | 78.2 |
| EgoCITE-GPT (ours) | 64.3 | 80.3 |
The EgoCITE-GPT ablation varies retrieval depth across one, three, and five rounds and reports retrieval hit rate and average QA accuracy on EgoLifeQA, with curation applied only to the three- and five-round variants.
Result: The five-round variant reaches a 49.6% hit rate and 63.8% average accuracy, improving over the one-round variant by 13.1 and 5.1 percentage points, respectively.
Insight: Most gains occur within three rounds, where hit rate rises by 12.1 points and accuracy by 3.9 points, showing that EgoCITE progressively refines retrieval and recovers additional relevant atomic memory indices.
| Metric | 1 roundNo curation | 3 roundsCurated | 5 roundsCurated |
|---|---|---|---|
| Retrieval and overall QA | |||
| Hit rate (%) | 36.5 | 48.6 | 49.6 |
| Average accuracy (%) | 58.7 | 62.6 | 63.8 |
| QA accuracy by category (%) | |||
| EntityLog | 61.9 | 66.1 | 67.4 |
| EventRecall | 54.7 | 58.6 | 59.5 |
| HabitInsight | 60.1 | 60.1 | 64.3 |
| RelationMap | 62.6 | 65.4 | 65.6 |
| TaskMaster | 61.0 | 71.0 | 71.4 |
Katrina and Lucia plan a flower pressing test in the courtyard.



I walk past the whiteboard and continue forward, heading outside. “Should we put this in a dry place?” Katrina asks. “It will press and dry the flowers faster,” she explains. I open the door and step into the courtyard. “Let's make another batch today,” Katrina says. “Remember we had a few extra flowers on the first day? Let's cut them and press them. See how it turns out. Because we have two days.” I approach the table. “I want to test it in advance,” Katrina adds. Lucia chimes in, “Let's see when to stop.”
Four complementary views: EgoIndex preserves entering the courtyard, flower-pressing suggestions, the flower-pressing activity, and its discussion.
Food preparation and a song discussion unfold in the same caption window.



I look around as I asks, “Where was I standing?” and then states, “I'm standing here.” Sweet potatoes are placed beside the item I am observing. I talk with Tasha and bring over a glove. Shure suggests, “It might be an intro,” to which I replies, “An intro.” Katrina enters while I am looking at the object, and she watches I look for songs next to Shure. I put the gloves on my hands, aligning them properly. Now wearing gloves, I observe that everyone is looking at the computer.
Four complementary views: EgoIndex preserves both the food-preparation behavior and the concurrent discussion about song structure.
The group discusses flat stones while practicing stone skipping.



Alice says, “If it's too flat, you can't grab it, hard to grab.” I reply, “Indeed, too flat is hard to grab,” as I throw a stone. I watch Tasha pick up stones. I adds, “Yeah, flat for stone skipping. Skipping, skipping stones. What's stone skipping? Really skipping stones.” I lose a stone and point in the direction where the stones are being thrown. I exclaims, “Look, I skipped it twice,” and laughs. I laugh happily, adjust my glasses, and place my right hand on the table before throwing another stone. I asks, “How many times, twice?” I lose another stone and put my leg down. Alice warns, “If you hit me, I'll tell you. I'll catch you with these glasses. Don't leave today. I'm telling you,” as I lower my head to pick up another stone.
Four complementary views: EgoIndex preserves individual stone throws, the stone-skipping practice activity, questions about stone skipping, and the group's discussion of which stones work.
Passing a phone around so each person can mark the timestamp.



“Come on, everyone mark it. Each person marks it,” I say. I put the phone in front of everyone and hand it to Katrina. “Pass it along and mark it, hey. You can see it, right?” I ask. Katrina hands the phone to Alice, who then passes it to Tasha. “Alright, all done marking,” I say. The phone is handed back to me. “Okay,” I say. I turn off my phone. “Turn it on,” I instruct myself, but instead put away my phone. I walk up to the small tripod. “So today we're just going to discuss a bit. Discuss this morning. What should we do on our last day? And then maybe... Hmm...” I say as I hold a morning meeting with everyone. I continue the meeting discussion with everyone.
[I | handed the phone to | Katrina]
[Katrina | handed the phone to | Alice]
[Alice | passed the phone to | Tasha]
[I | hold a morning meeting with | everyone]
Issue: The handoff triples omit the purpose of marking a timestamp.
Jake handed his phone to Katrina to pass along for marking.
Jake received his phone back and turned it off before putting it away.
Jake walked to a small tripod and led a morning meeting about what to do on the last day.
Issue: Individual observations compound multiple actions, reducing the precision of similarity-based retrieval.
UtteranceI tell everyone to mark the timestamp on my phone.
ActionI hand my phone to Katrina.
ActionKatrina hands my phone to Alice.
ActionAlice hands my phone to Tasha.
Advantage: EgoScheme preserves the marking intent while separating each handoff into its own searchable atomic memory index.
Jake puts boxes down before the group looks for a pen.



I hold the box to my chest and walk forward, turning right before placing it in the corner. “This thing, you can put these down,” I says. Shure adds, “Let's assemble it later, let's write first.” I turn the box with both hands as I continue walking. “It's not a big problem, right?” I asks. “Not a big problem,” he confirms, while Shure repeats, “Not a big problem, not a big problem. It won't affect my calligraphy.” I place the box in the corner and turn around, walking toward the dining table. “Where's the pen?” I asks. Shure echoes, “Where's the pen?” and remarks, “Good question,” which I agrees with. I walk left into my room and look at the table, then glance back at the bed. “Ask him, ask AI, where's the pen?” Shure suggests. “Where's the pen, where's the pen?” I repeats. Lucia chimes in, “Looking for my iPhone?” and laughs. I turn right and walk out of the room, continuing forward as I asks again, “Yeah, where's the pen?” and Lucia starts, “Looking for my...”
[I | place | box]
[Shure | adds | Let's assemble it later]
[I | asks | Where's the pen?]
[Lucia | starts | Looking for my...]
Issue: The triples preserve incomplete and context-dependent speech such as “Let's assemble it” and “Looking for my...”
Jake carried a box against his chest, placed it in the corner, and said those things could be put down.
Jake went into his room, looked around, and came back out still asking for the pen.
Shure suggested assembling the item later and writing first.
Issue: “Those things” and “the item” remain generic, while several distinct behaviors are merged into single observations.
ActionI place the box in the corner of the wall.
UtteranceI tell Shure these boxes can be put down.
UtteranceShure suggests assembling the box later.
UtteranceLucia asks me whether I am looking for Lucia's iPhone.
Advantage: EgoScheme replaces generic references with named people and objects, then decouples the actions and utterances.
Katrina asks whether all packages are outside; Jake answers “Yes.”



I stand still, turn my head to the left, and walk forward into the entryway. I continue walking through the hall while Katrina asks, “Are all the packages outside?” I reply, “Yes.” I adds, “Today, a few more arrived. But the housekeeper will still come to deliver them.” I pick up a pole and speak while holding it.
[Katrina | asks | Are all the packages outside?]
[I | reply | Yes]
[housekeeper | will come to deliver | packages]
Issue: “[I | reply | Yes]” is not self-contained; its meaning is lost when retrieved without the preceding question.
Jake replied, “Yes.” when asked if all packages were outside.
Jake said, “Today, a few more arrived. But the housekeeper will still come to deliver them.”
Jake picked up a pole and spoke while holding it.
Issue: The observation preserves the elliptical answer verbatim instead of expressing the confirmed proposition directly.
UtteranceKatrina asks about the packages outside.
UtteranceI reply that the packages are outside.
UtteranceI say that a few more packages arrived today.
UtteranceI add that the housekeeper will deliver more packages later.
Advantage: EgoScheme resolves “Yes” into a complete proposition and separates the exchange into searchable atomic memory indices.
QuestionWhat was I doing the first time Tasha made cupcakes?
ActionI fry eggsDAY1 19000000 to DAY1 20300000Why the rank changes: The target lies inside the requested time window and keeps full temporal relevance, while the DAY4 semantic match is downweighted from 0.758 to 0.588 and falls from #1 to #5; the target consequently enters the action view's top-20 retrieval budget.
QuestionWho was I with the last time I stood on the escalator in the supermarket?
Activitygo supermarket escalatorDAY3 17050000 to DAY3 17200000Why the rank changes: The answer-bearing group activity lies inside the drafting agent's “last time” window and keeps full temporal relevance; downweighting out-of-window supermarket matches moves it 17 places into the activity view's top-10 retrieval budget.
QuestionWhen was the last time Wonky's Milk appeared in front of me?
Actionmilk cartonDAY1 11000000 to DAY1 16000000Why the rank changes: The target lies inside the DAY1 time window and keeps full temporal relevance, while the closest DAY3 milk match is downweighted from 0.585 to 0.440 and falls from #1 to #11; the target consequently enters the action view's top-20 retrieval budget.
Showing First cupcake event. The target-overlapping index moves from rank 24 to rank 4.
BibTeX entry for the arXiv preprint.
@misc{zhang2026egocitecontextaugmentedindexingtimeaware,
title={EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory},
author={Le Zhang and Hao Chen and Vlad Roznyatovskiy and Jianzhong Zhang and Ke Sun},
year={2026},
eprint={2608.12627},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.12627},
}