Learning from human-computer interaction histories to predict when AI agents need intervention
Multimodal Agent Interface Interactions
Train agents to recognize when they're about to make mistakes by learning from patterns in successful vs. failed interaction sequences.
Context: Prior work (e.g., memory-augmented LLM agents, generative human-behavior simulators) shows that agents can integrate past interactions into behavior. However, they lack full capture of multimodal human-agent interaction traces, correction loops, and lightweight adaptation in UI contexts. Chain of Experience logs UI state (DOM structure, screenshots, interactive elements), records user interventions and corrections, and training experience encoders for agents to learn from how users interact with them.
python agent/run_agent.pyThis runs the agent with:
- LangGraph + Async Playwright
- MP4 video recording with visible cursor
- Intervention logging in SQLite
- ReAct pattern (Reasoning → Action → Observation)
- Observation (interactive elements + accessibility tree)
- Collect interaction data - Record agent browsing sessions with interventions
- Learn patterns - Train models to recognize pre-intervention states
- Predict interventions - Detect when agent is about to need help
- Improve autonomy - Agents learn when to ask vs. when to continue
Instead of just recording what the agent did, we record when a user would have intervened, creating training data for learning failure patterns.
- LangGraph framework
- Async Playwright
- ReAct pattern - Think -> Act -> Observe loop
- Task tracking - SQLite database logging
- Visible cursor that persists across pages
- Smooth scrolling and element highlighting
- WebM -> MP4 conversion
- Intervention logging
- Page state at each step
- Action history tracking
- Success/failure metrics
Running the agent produces:
- MP4 videos (
data/videos/) - For demos and papers - Task logs (
data/tasks.db) - For training models - Statistics - Intervention patterns and success rates
# Install dependencies
poetry install
# Install Playwright browsers
poetry run playwright install chromium
# Install ffmpeg for video conversion
brew install ffmpeg # macOS
# or: apt-get install ffmpeg # LinuxSee agent/README.md for full documentation.
Try these to see what the agent can do:
E-commerce:
- "Find the best wireless mouse under $50"
- "Compare noise-cancelling headphones on Amazon"
- "Look for a mirrorless camera with good reviews"
Research:
- "Find the latest Python release notes"
- "Search for papers about LLMs on arXiv"
- "Look up Playwright documentation"
Information:
- "Get the weather forecast for New York"
- "Find top-rated pizza restaurants in Chicago"
- "Search for news about AI developments"
Navigation:
- "Find the PyTorch installation guide"
- "Look up LangGraph tutorials"
- "Search for async Playwright examples"
import asyncio
from agent.run_agent import run_consolidated_agent
async def main():
result = await run_consolidated_agent(
url="https://en.wikipedia.org",
goal="Find information about Python programming language",
max_steps=20,
headless=False
)
print(f"Status: {result['status']}")
print(f"Video: {result['video_path']}")
asyncio.run(main())Run the agent on diverse tasks to build a dataset of intervention patterns:
# Initialize database
poetry run python scripts/init_db.py
# Run agent on various tasks
poetry run python agent/run_agent.py "Find the best wireless mouse"
poetry run python agent/run_agent.py "Search for Python documentation"
# ... run 20-50 diverse tasks
# Export training data
poetry run python scripts/export_training_data.pyThis creates:
- Intervention states - When agent gets stuck or needs help
- Success states - When agent completes tasks smoothly
- Video recordings - Visual context for each state
- Page states - DOM structure and interaction history
Train the intervention prediction model using contrastive learning:
# Train (2-phase training)
poetry run python train.py --epochs-contrastive 20 --epochs-predictor 10
# Evaluate
poetry run python evaluate.pyModel Architecture:
-
Experience Encoder - Multimodal encoder combining:
- Visual (CLIP on screenshots/video frames)
- Actions (transformer over action sequences)
- Page state (DOM structure, accessibility tree)
-
Contrastive Learning - Learn to distinguish:
- Similar states (both successful or both needing intervention)
- Dissimilar states (successful vs. intervention needed)
-
Intervention Predictor - Binary classifier predicting when agent needs help
The evaluation script computes:
- Accuracy - Overall prediction correctness
- Precision/Recall - Intervention prediction quality
- F1 Score - Harmonic mean of precision/recall
- ROC AUC - Model discrimination ability
- Confusion Matrix - Breakdown of prediction types
Results are saved to data/visualizations/ with plots and metrics.
- LangGraph: https://github.com/langchain-ai/langgraph
- Playwright: https://playwright.dev
- ReAct Pattern: https://arxiv.org/abs/2210.03629
- talk2browser: https://github.com/talk2silicon/talk2browser
This project is licensed under the MIT License - see the LICENSE file for details.
Authors: Danylo Voloshyn, Diego Abarcar-Calugay, Joanne Lee, Om Arya, Yonatan Tussa