Skip to content

About

PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks

Topics

Resources

Stars

17 stars

Watchers

0 watching

Forks

Repository files navigation

PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks

arXiv alphaXiv Dataset

Bowen Jiang, Yuan Yuan, Zhuoqun Hao, Yuchen Liu, Maohao Shen, Sihao Chen, Gregory Wornell, Chris Callison-Burch, Lyle Ungar, Dan Roth, Qi Guo, Xiangjun Fan, Camillo J. Taylor, Hanchao Yu

A collaboration between Meta Meta Recommendation Systems, UPenn University of Pennsylvania, and MIT MIT.

PersonaMem-v3

Third release in the PersonaMem series:

  • PersonaMem (v1) — [COLM 2025] Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale · code · paper · data
  • PersonaMem-v2 — Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory · code · paper · data

What's new in v3

Dimension PersonaMem-v1 PersonaMem-v2 PersonaMem-v3
Data source 20 fully synthetic users 1000 fully synthetic users with more comprehensive personas 200 anonymized real-world users with 4,000,000 engagement histories
Explicit vs. implicit Explicit user preferences Implicit user preferences Around 95% implicit user behavior signals
Scenarios Chatbot conversations Chatbot conversations Omni-platform, including chatbot, social media recommendation, agentic tasks, and proactiveness
Restraint Personalization Personalization Personalization and over-personalization
User privacy No mentioning of user private information Including personally identifiable information and user-initiated ask-to-forget scenarios Including psychology-anchored hidden persona and socially inappropriate scenarios
Dynamics Fully synthesized preference updates Fully synthesized preference updates Reinforced, emerging, diminishing, bursting, and varied attention shifts from the real world

Setup

pip install -r requirements.txt
cp .env.example .env        # fill in Azure OpenAI / Gemini credentials
# download the full source engagement data (facebook/gistbench) and convert it to the input CSV:
python scripts/download_gistbench.py                # → data/gistbench_input.csv

Source data comes from facebook/gistbench; pre-built personas are distributed via the PersonaMem-v3 dataset. Neither is tracked in git — but a ready-made 10-user sample input ships with the repo at data/gistbench_sample_10users.csv for a quick start.

1. Build personas, user histories, their queries

One persona, end to end (full generation pipeline → backend/115/ with profile.json, five app histories, calendar.json, persona.html):

python scripts/run_persona_pipeline.py --user_id 115 --input_csv data/gistbench_input.csv --verbose
python scripts/prepare_eval_data.py --user_id 115        # → backend/115/test.json (the eval queries)

Multiple personas:

bash scripts/run_persona_batch.sh                         # every user in the input CSV; resumable
NUM_USERS=25 bash scripts/run_persona_batch.sh            # only the first 25 user ids in the CSV
USERS="17 18 115" bash scripts/run_persona_batch.sh       # explicit persona ids
# other knobs: INPUT_CSV=path/to/input.csv  CONCURRENCY=3 (personas generated simultaneously)
python scripts/prepare_eval_data.py --all --parallel 4    # queries for every generated persona

2. Run evaluations and show results

All artifacts are plain JSON/CSV/HTML on disk:

  • Personas and their data — backend/{uid}/: profile.json (persona definition), instagram.json / facebook.json / threads.json / chatbot.json / ai_studio.json (time-sorted interaction-event histories per app), calendar.json (calendar modification stream), test.json (the eval queries), persona.html (self-contained human-readable review page).
  • Eval runs — results/{mode}/{uid}/: results.csv (one row per query: query_id, seq, user_id, task_type, ts, metrics_json, status, duration_ms, error, agent_response), writes.jsonl (agentic write actions), summary.json.
  • Aggregates — results/aggregate/ (per-mode CSV/JSON summaries); final comparison tables at results/aggregate/html/results_tables.html.

Single persona, single mode:

python evaluation/run_eval.py --user_id 115 --mode llm_longctx \
    --model gpt-5.5 --judge_model gpt-5.5 --run_dir results/llm_longctx_gpt5.5/115

--mode ∈ llm_longctx (long-context baseline) · llm_memory (textual memory) · mem0 (mem0 memory) · claude_code (Claude Code agent over time-masked filesystem snapshots) · codex (Codex CLI agent over the same snapshots). The judge is always gpt-5.5.

Full matrix over a cohort of personas and modes:

scripts/run_eval_matrix.sh --personas "101 102 103" --modes "llm_longctx llm_memory mem0"

Aggregate and render the summary tables:

python scripts/aggregate_eval.py --results_root results   # → results/aggregate/ (CSV/JSON summaries)

Note for agents (Claude Code / Codex): all commands above are directly runnable from the repo root; evaluation runs make real LLM API calls, so confirm with the user before launching them.

License

Code is released under the MIT License. The sample input data and all personas derived from facebook/gistbench inherit its CC-BY-NC-4.0 license (attribution, non-commercial).

Citations

If you find our work helpful, please consider cite them. Thank you!

@article{jiang2026personamem,
  title={PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks},
  author={Jiang, Bowen and Yuan, Yuan and Hao, Zhuoqun and Liu, Yuchen and Shen, Maohao and Chen, Sihao and Wornell, Gregory and Callison-Burch, Chris and Ungar, Lyle and Roth, Dan and others},
  journal={arXiv preprint arXiv:2608.21381},
  year={2026}
}

@article{jiang2025personamem2,
  title={PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory},
  author={Jiang, Bowen and Yuan, Yuan and Shen, Maohao and Hao, Zhuoqun and Xu, Zhangchen and Chen, Zichen and Liu, Ziyi and Vijjini, Anvesh Rao and He, Jiashu and Yu, Hanchao and Poovendran, Radha and Wornell, Gregory and Ungar, Lyle and Roth, Dan and Chen, Sihao and Taylor, Camillo Jose},
  journal={arXiv preprint arXiv:2512.06688},
  year={2025}
}

@article{jiang2025know,
  title={Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale},
  author={Jiang, Bowen and Hao, Zhuoqun and Cho, Young-Min and Li, Bryan and Yuan, Yuan and Chen, Sihao and Ungar, Lyle and Taylor, Camillo J and Roth, Dan},
  journal={arXiv preprint arXiv:2504.14225},
  year={2025}
}

About

PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks

Topics

Resources

Stars

17 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages