
Applied AI Researcher
I build agents
that remember.
Research and systems for agent memory, attention, and reliable computer use. First author of Proced-Mem, a benchmark for procedural memory retrieval in language agents (ICLR 2026 MemAgents Workshop). Currently building CASE.
Selected Work
Research
MemAgents Workshop, ICLR 2026
Proced-Mem: Benchmarking Procedural Memory Retrieval in Language Agents Across Domains
Ishant Kohar and Aswanth Krishnan
The first benchmark to isolate procedural memory retrieval from task execution, spanning text-based household tasks (ALFWorld) and real computer-use environments (OSWorld). Evaluates seven retrieval methods across text, visual, and lexical modalities with an LLM-as-judge protocol validated against human annotations.
Embedding-based retrieval degrades 30 to 42% on novel contexts, a generalization cliff: encoders treat procedures as unordered bags of tokens and discard the temporal structure that transfer depends on.