Skip to content

Latest commit

 

History

258 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Awesome LLM-Based Human-Agent Collaboration and Interaction Systems

🎉 Our survey has been accepted to ACL 2026!

ACL 2026 Awesome arXiv Maintenance Contribution Welcome Last Commit GitHub stars

Oryx Video-ChatGPT

image

Welcome to Awesome-Human-Agent-Collaboration-Interaction-Systems! 🚀 This is the repo for our Survey on LLM-Based Human-Agent Collaboration and Interaction Systems, accepted to ACL 2026.

🌟 Introduction

Recent advances in large language models (LLMs) have sparked growing interest in building fully autonomous agents. However, fully autonomous LLM-based agents still face significant challenges, including (1) limited reliability due to hallucinations, (2) difficulty in handling complex tasks, and (3) substantial safety and ethical risks, all of which limit their feasibility and trustworthiness in real-world applications.

LLM-based human-agent collaboration systems are interactive frameworks where humans actively provide (1) additional information, (2) feedback, or (3) control during interaction with LLM-powered agents to enhance system performance, reliability, and safety. These human-agent collaboration systems enable humans and LLM-based agents to collaborate effectively by leveraging their complementary strengths. For a detailed introduction, please refer to our survey paper: LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey.

Our goal with this project is to build an exhaustive collection of awesome resources relevant to LLM-Based Human-Agent/AI Collaboration and Interaction Systems, encompassing papers, repositories, and more to foster further research and innovation in this rapidly evolving interdisciplinary field of human-ai collaboration. 🤗 Contributions are welcome! 🤗 If you have recommended papers, resources or suggestions, please submit pull requests, open issues or contact us. We will keep updating our repo & survey paper.

📄 Contents

📄 Latest Research Papers

(©️click here back to table of contents👆🏻)

🤗 Contributions are welcome! If you have recommended papers and resources, please submit pull requests or open issues.

📚 Applications, Datasets & Benchmarks

(©️click here back to table of contents👆🏻)

💻 Web Navigation & Computer Use

👨🏻‍💻 Software Engineering, Coding

🤖 Embodied AI, Robotics

💬 Conversation System

📊 Data Science, Scientific Discovery

🎮 Gaming

💰 Finance

🏥 Healthcare, Medicine

🧩 General-Purpose Assistants, Cross-Domain

🛍️ Retail, Telecom

🛩️ Travel

✍️ Writing

🔍 Taxonomy

(©️click here back to table of contents👆🏻)

For a detailed introduction of the taxonomy, please refer to Section 3 in our survey paper: LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey.

image


🤝 Human Feedback

(©️click here back to table of contents👆🏻)

Human Feedback can occur during different phases in various types and granularities. In the following table, we summarize different dimensions of Human Feedback in LLM-based human-agent systems, including feedback type, granularity, and phase. For each dimension, a summary, key characteristics, and example works are provided for comparison. More details are in Section 3.2 of our survey paper.

image

Title Date & Code Feedback Type Feedback Subtype Feedback Granularity Feedback Phase
Just A Rather Very Intelligent Spoken Agent (JarvisBench) 2026/07 Guidance, Corrective Refinement, Critique Segment During Task
LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability 2026/07 GitHub stars Guidance, Evaluative Critique, Binary Assessment Segment During Task
HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation 2026/07 Guidance, Evaluative, Corrective, Implicit Critique, Refinement, Binary Assessment, Human Control Holistic, Segment Initial Setup, During Task, Post Task
Uncertainty Decomposition for Clarification Seeking in LLM Agents 2026/06 Guidance Demonstration Segment During Task
Learning User Simulators with Turing Rewards 2026/06 Implicit User Action Segment During Task
Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents (TRACE) 2026/06 GitHub stars Corrective, Guidance Refinement, Critique Segment During Task, Post Task
Re-Centering Humans in LLM Personalization 2026/06 Evaluative, Implicit Preference Ranking, Scalar Rating, User Action Holistic, Segment Post Task
CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive Engagement 2026/06 Implicit, Guidance User Action, Critique Segment During Task
Human Oversight of Agentic Systems in Practice: Examining the Oversight Work, Challenges, and Heuristics of Developers Using Software Agents 2026/06 Evaluative, Corrective, Guidance Binary Assessment, Refinement, Critique Holistic, Segment Initial Setup, During Task, Post Task
Uncertainty-Aware Clarification in LLM Agents with Information Gain 2026/06 Guidance Demonstration Segment During Task
Not All Uncertainty Is Equal: How Uncertainty Granularity Shapes Human Verification in LLM-Assisted Decision Making 2026/05 Evaluative, Implicit Binary Assessment, User Action Segment, Holistic During Task
VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions 2026/05 GitHub stars Implicit, Guidance User Action, Demonstration Segment Initial Setup, During Task
Reinforcing Human Behavior Simulation via Verbal Feedback (DITTO & SOUL) 2026/05 Guidance, Corrective Critique, Refinement Holistic, Segment Post Task
ECHO: Explainable Co-editing with Human-in-the-loop Operations for Presentation Refinement 2026/05 Corrective, Guidance Refinement, Demonstration Segment During Task
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators 2026/05 Guidance, Evaluative Critique, Binary Assessment Segment Initial Setup, During Task
SalesSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators 2026/05 Implicit, Guidance User Action, Critique Segment During Task
A Decoupled Human-in-the-Loop System for Controlled Autonomy in Agentic Workflows 2026/04 Evaluative, Corrective, Implicit Binary Assessment, Refinement, Human Control Segment Initial Setup, During Task
CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks 2026/04 Guidance, Evaluative, Implicit Demonstration, Scalar Rating, User Action Holistic, Segment During Task, Post Task
When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation (InterruptBench) 2026/04 GitHub stars Corrective, Guidance Refinement, Critique Segment During Task
Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents 2026/03 Guidance Demonstration, Critique Segment During Task
User Preference Modeling for Conversational LLM Agents: Weak Rewards from Retrieval-Augmented Interaction (VARS) 2026/03 GitHub stars Evaluative, Implicit Scalar Rating, User Action Segment During Task, Post Task
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks 2026/03 GitHub stars Implicit, Guidance User Action, Demonstration Segment Initial Setup, During Task
AgentDS: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science 2026/03 Guidance, Corrective, Evaluative Demonstration, Refinement, Scalar Rating Holistic, Segment Initial Setup, During Task, Post Task
ViviDoc: Generating Interactive Documents through Human-Agent Collaboration 2026/03 GitHub stars Corrective, Guidance Refinement, Demonstration Segment During Task
InfoPO: Information-Driven Policy Optimization for User-Centric Agents 2026/02 GitHub stars Guidance, Implicit Demonstration, User Action Segment During Task
Modeling Distinct Human Interaction in Web Agents (CowCorpus) 2026/02 Implicit, Corrective User Action, Human Control, Refinement Segment During Task
Overseeing Agents Without Constant Oversight: Challenges and Opportunities 2026/02 Evaluative, Corrective Binary Assessment, Critique Segment, Holistic During Task, Post Task
Learning Personalized Agents from Human Feedback 2026/02 GitHub stars Guidance, Corrective, Evaluative Critique, Refinement, Preference Ranking Segment Initial Setup, During Task, Post Task
Value of Information: A Framework for Human-Agent Communication 2026/01 Guidance Demonstration Segment During Task
Progressive Ideation using an Agentic AI Framework for Human-AI Co-Creation (MIDAS) 2026/01 Evaluative, Guidance Preference Ranking, Critique Segment During Task
From Correctness to Collaboration: Toward a Human-Centered Framework for Evaluating AI Agent Behavior in Software Engineering 2025/12 Guidance, Evaluative Critique, Preference Ranking Segment Initial Setup, During Task
CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding 2025/11 Guidance, Corrective Refinements, Critique Segment, Holistic During Task
Training Proactive and Personalized LLM Agents 2025/11 GitHub stars Guidance, Evaluative Scalar Rating, Refinements Segment, Holistic During Task
Training LLM Agents to Empower Humans 2025/10 GitHub stars Implicit, Corrective User Action, Refinements Segment During Task
How can we assess human-agent interactions? Case studies in software agent design 2025/10 Evaluative Scalar Rating Segment During Task
RECODE-H: A Benchmark for Research Code Development with Interactive Human Feedback 2025/10 GitHub stars Guidance, Corrective Refinements, Critique, Demonstration Segment During Task
UserRL: Training Proactive User-Centric Agent via Reinforcement Learning 2025/09 GitHub stars Evaluative, Guidance, Implicit Scalar Rating, Refinements, Critique Segment During Task
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use 2025/08 GitHub stars Evaluative, Guidance Binary Assessment, Refinement Holistic Initial Setup, During Task
Magentic-UI: Towards Human-in-the-loop Agentic Systems 2025/07 GitHub stars Evaluative, Corrective, Guidance, Implicit Binary Assessment, Refinement, Critique, User Action Segment During Task, Post Task
UserBench: An Interactive Gym Environment for User-Centric Agents 2025/07 GitHub stars Implicit, Guidance User Action, Refinement Segment Initial Setup, During Task
Enabling Self-Improving Agents to Learn at Test Time With Human-In-The-Loop Guidance 2025/07 GitHub stars Corrective, Guidance Refinement, Critique Segment During Task
τ2-Bench: Evaluating Conversational Agents in a Dual-Control Environment 2025/06 GitHub stars Evaluative, Implicit Binary Assessment, User Action Segment, Holistic Initial Setup, During Task
Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training 2025/05 Corrective Refinement Segment During Task
Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild 2025/05 GitHub stars Corrective, Guidance, Evaluative Refinement, Binary Assessment, Demonstration Segment, Holistic During Task
XtraGPT: LLMs for Human-AI Collaboration on Controllable Academic Paper Revision 2025/05 GitHub stars Guidance, Evaluative Critique, Preference Ranking Segment Initial Setup, Post Task
MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft 2025/04 GitHub stars Evaluative Scalar Rating Holistic Post Task
Experimental Exploration: Investigating Cooperative Interaction Behavior Between Humans and Large Language Model Agents 2025/03 Implicit User Action Segment During Task
SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks 2025/03 GitHub stars Corrective, Implicit Refinement, User Action Segment During Task
FinArena: A Human-Agent Collaboration Framework for Financial Market Analysis and Forecasting 2025/03 Guidance Demonstration Segment During Task
Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI Collaboration 2025/02 GitHub stars Guidance Critique Holistic During Task
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments 2025/02 GitHub stars Guidance Demonstration, Critique Segment During Task
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation 2025/01 Corrective, Implicit User Action, Refinement Segment During Task
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration (Co-Gym) 2024/12 GitHub stars Corrective Refinement Segment During Task
Towards Modeling Human-Agentic Collaborative Workflows: A BPMN Extension 2024/12 GitHub stars Guidance Demonstration Holistic Initial Setup, During Task, Post Task
To Help or Not to Help: LLM-based Attentive Support for Human-Robot Group Interactions 2024/10 GitHub stars Implicit, Guidance Demonstration, User Action Holistic Initial Setup, During Task
PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks 2024/10 GitHub stars Corrective, Guidance Refinement, Critique Holistic, Segment During Task, Post Task
AssistantX: An LLM-Powered Proactive Assistant in Collaborative Human-Populated Environment 2024/09 GitHub stars Implicit Human Control Holistic Initial Setup
Mutual Theory of Mind in Human-AI Collaboration: An Empirical Study with LLM-driven AI Agents in a Real-time Shared Workspace Task 2024/09 Corrective Refinement Segment During Task
Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent Conversations 2024/08 GitHub stars Guidance Demonstration Segment During Task
Human-LLM collaboration in generative design for customization 2024/07 Guidance, Evaluative Demonstration, Binary Assessment, Preference Ranking Holistic, Segment Initial Setup, Post Task
WebCanvas: Benchmarking Web Agents in Online Environments 2024/06 Evaluative Scalar Rating Holistic Post Task
Enhancing Human-Robot Collaborative Assembly in Manufacturing Systems Using Large Language Models 2024/06 Corrective, Guidance Demonstration, Refinement, Critique Segment Initial Setup, During Task
Enhancing the LLM-Based Robot Manipulation Through Human-Robot Collaboration 2024/06 Corrective Refinement Holistic During Task
DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning 2024/06 GitHub stars Evaluative Binary Assessment Holistic During Task, Post Task
REVECA: Adaptive Planning and Trajectory-based Validation in Cooperative Language Agents using Information Relevance and Relative Proximity 2024/05 Implicit Human Control Segment During Task
Autonomous Evaluation and Refinement of Digital Agents 2024/04 GitHub stars Evaluative Binary Assessment Holistic Post Task
A Human-Computer Collaborative Tool for Training a Single Large Language Model Agent into a Network through Few Examples 2024/04 Corrective, Guidance Demonstration, Refinement Segment During Task
An LLM-based approach for Enabling Seamless Human-Robot Collaboration in Assembly 2024/04 Guidance, Corrective Demonstration, Refinement Segment During Task
AgentCoord: Visually Exploring Coordination Strategy for LLM-based Multi-Agent Collaboration 2024/04 GitHub stars Guidance, Corrective Demonstration, Refinement Segment Initial Setup
PDFChatAnnotator: A Human-LLM Collaborative Multi-Modal Data Annotation Tool for PDF-Format Catalogs 2024/04 Corrective, Guidance Demonstration, Refinement Segment During Task
Embodied LLM Agents Learn to Cooperate in Organized Teams 2024/03 GitHub stars Guidance Critique Holistic During Task
Large Language Model-based Human-Agent Collaboration for Complex Task Solving 2024/02 GitHub stars Implicit Human Control Segment During Task
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue 2024/02 GitHub stars Guidance Demonstration Holistic, Segment During Task
Ask-before-Plan: Proactive Language Agents for Real-World Planning 2024/01 GitHub stars Guidance Demonstration, Critique Segment During Task
A2C: A Modular Multi-stage Collaborative Decision Framework for Human-AI Teams 2024/01 Guidance, Implicit, Corrective, Evaluative Refinement, Binary Assessment, Critique, Human Control Holistic, Segment During Task
LLM-Powered Hierarchical Language Agent for Real-time Human-AI Coordination 2023/12 GitHub stars Guidance, Corrective Demonstration, Refinement Segment During Task
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents 2023/10 Evaluative, Implicit Scalar Rating, User Action Holistic, Segment During Task, Post Task
MindAgent: Emergent Gaming Interaction 2023/09 GitHub stars Corrective Refinement Segment During Task
Drive As You Speak: Enabling Human-Like Interaction With Large Language Models in Autonomous Vehicles 2023/09 Guidance Demonstration Segment During Task
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback 2023/09 Evaluative Binary Assessment Holistic During Task
MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework 2023/08 GitHub stars Evaluative, Guidance Binary Assessment Holistic Initial Setup, Post Task
LLM-Based Human-Robot Collaboration Framework for Manipulation Tasks 2023/08 Corrective, Guidance Demonstration, Refinement Segment During Task
Embodied Task Planning with Large Language Models 2023/07 GitHub stars Guidance, Evaluative Demonstration, Binary Assessment Holistic, Segment Initial Setup, Post Task
Building cooperative embodied agents modularly with large language models 2023/07 Evaluative Scalar Rating Holistic Post Task
Improved Trust in Human-Robot Collaboration With ChatGPT 2023/06 Guidance Demonstration, Critique Segment During Task
Investigating Agency of LLMs in Human-AI Collaboration Tasks 2023/05 Guidance Demonstration, Critique Segment During Task
Improving grounded language understanding in a collaborative environment by interacting with agents through help feedback 2023/04 Corrective, Guidance Demonstration, Refinement Holistic During Task
PaLM-E: An Embodied Multimodal Language Model 2023/03 Guidance, Implicit Demonstration, User Action Segment During Task
Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts 2021/10 Corrective Refinement Segment During Task

🔄 Interaction

(©️click here back to table of contents👆🏻)

Title Date & Code Interaction Types Interaction Variant
Just A Rather Very Intelligent Spoken Agent (JarvisBench) 2026/07 Collaboration Supervision, Cooperation
LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability 2026/07 GitHub stars Collaboration Cooperation, Coordination
HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation 2026/07 Collaboration Supervision, Delegation, Coordination
Uncertainty Decomposition for Clarification Seeking in LLM Agents 2026/06 Collaboration Cooperation
Learning User Simulators with Turing Rewards 2026/06 Collaboration Cooperation
Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents (TRACE) 2026/06 GitHub stars Collaboration Supervision, Cooperation
Re-Centering Humans in LLM Personalization 2026/06 Collaboration Cooperation
CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive Engagement 2026/06 Collaboration Cooperation, Coordination
Human Oversight of Agentic Systems in Practice: Examining the Oversight Work, Challenges, and Heuristics of Developers Using Software Agents 2026/06 Collaboration Supervision, Delegation
Uncertainty-Aware Clarification in LLM Agents with Information Gain 2026/06 Collaboration Cooperation
Not All Uncertainty Is Equal: How Uncertainty Granularity Shapes Human Verification in LLM-Assisted Decision Making 2026/05 Collaboration Supervision
VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions 2026/05 GitHub stars Collaboration Cooperation, Delegation
Reinforcing Human Behavior Simulation via Verbal Feedback (DITTO & SOUL) 2026/05 Collaboration Cooperation
ECHO: Explainable Co-editing with Human-in-the-loop Operations for Presentation Refinement 2026/05 Collaboration Supervision, Cooperation
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators 2026/05 Collaboration Coordination
SalesSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators 2026/05 Collaboration Cooperation
A Decoupled Human-in-the-Loop System for Controlled Autonomy in Agentic Workflows 2026/04 Collaboration Supervision, Delegation
CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks 2026/04 Collaboration Cooperation, Delegation
When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation (InterruptBench) 2026/04 GitHub stars Collaboration Supervision, Cooperation
Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents 2026/03 Collaboration Cooperation
User Preference Modeling for Conversational LLM Agents: Weak Rewards from Retrieval-Augmented Interaction (VARS) 2026/03 GitHub stars Collaboration Cooperation
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks 2026/03 GitHub stars Collaboration Delegation
AgentDS: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science 2026/03 Collaboration Cooperation, Delegation
ViviDoc: Generating Interactive Documents through Human-Agent Collaboration 2026/03 GitHub stars Collaboration Cooperation
InfoPO: Information-Driven Policy Optimization for User-Centric Agents 2026/02 GitHub stars Collaboration Cooperation
Modeling Distinct Human Interaction in Web Agents (CowCorpus) 2026/02 Collaboration Supervision, Delegation, Coordination
Overseeing Agents Without Constant Oversight: Challenges and Opportunities 2026/02 Collaboration Supervision
Learning Personalized Agents from Human Feedback 2026/02 GitHub stars Collaboration Cooperation
Value of Information: A Framework for Human-Agent Communication 2026/01 Collaboration Cooperation
Progressive Ideation using an Agentic AI Framework for Human-AI Co-Creation (MIDAS) 2026/01 Collaboration Cooperation
From Correctness to Collaboration: Toward a Human-Centered Framework for Evaluating AI Agent Behavior in Software Engineering 2025/12 Collaboration Supervision, Cooperation
CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding 2025/11 Collaboration Delegation, Cooperation
Training Proactive and Personalized LLM Agents 2025/11 GitHub stars Collaboration Cooperation
Training LLM Agents to Empower Humans 2025/10 GitHub stars Collaboration Cooperation
How can we assess human-agent interactions? Case studies in software agent design 2025/10 Collaboration Delegation, Supervision
RECODE-H: A Benchmark for Research Code Development with Interactive Human Feedback 2025/10 GitHub stars Collaboration Supervision, Cooperation
UserRL: Training Proactive User-Centric Agent via Reinforcement Learning 2025/09 GitHub stars Collaboration Supervision, Cooperation
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use 2025/08 GitHub stars Collaboration Delegation
Magentic-UI: Towards Human-in-the-loop Agentic Systems 2025/07 GitHub stars Collaboration Cooperation, Coordination
UserBench: An Interactive Gym Environment for User-Centric Agents 2025/07 GitHub stars Collaboration Cooperation
Enabling Self-Improving Agents to Learn at Test Time With Human-In-The-Loop Guidance 2025/07 GitHub stars Collaboration Supervision
τ2-Bench: Evaluating Conversational Agents in a Dual-Control Environment 2025/06 GitHub stars Collaboration Cooperation, Coordination
Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training 2025/05 Collaboration Cooperation
Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild 2025/05 GitHub stars Collaboration Supervision, Cooperation
XtraGPT: LLMs for Human-AI Collaboration on Controllable Academic Paper Revision 2025/05 GitHub stars Collaboration Delegation, Supervision
MineWorld: A Real-Time and Open-Source Interactive World Model on Minecraft 2025/04 GitHub stars Collaboration Delegation
Experimental Exploration: Investigating Cooperative Interaction Behavior Between Humans and Large Language Model Agents 2025/03 Coopetition -
FinArena: A Human-Agent Collaboration Framework for Financial Market Analysis and Forecasting 2025/03 Collaboration Delegation
SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks 2025/03 GitHub stars Collaboration Delegation
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments 2025/02 GitHub stars Collaboration Supervision, Delegation
Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI Collaboration 2025/02 GitHub stars Collaboration -
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation 2025/01 Collaboration Supervision, Delegation, Coordination
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration (Co-Gym) 2024/12 GitHub stars Collaboration Supervision, Delegation
Towards Modeling Human-Agentic Collaborative Workflows: A BPMN Extension 2024/12 GitHub stars Collaboration Coordination
To Help or Not to Help: LLM-based Attentive Support for Human-Robot Group Interactions 2024/10 GitHub stars Collaboration Coordination
PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-Agent Tasks 2024/10 GitHub stars Collaboration Coordination, Cooperation
AssistantX: An Proactive Assistant in Collaborative Human-Populated Environment 2024/09 GitHub stars Collaboration Delegation
Mutual Theory of Mind in Human-AI Collaboration: An Empirical Study with LLM-driven AI Agents in a Real-time Shared Workspace Task 2024/09 Collaboration Coordination
Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent Conversations 2024/08 GitHub stars Collaboration Delegation
Human-LLM Collaboration in Generative Design for Customization 2024/07 Collaboration Delegation
DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning 2024/06 GitHub stars Collaboration Delegation
Enhancing Human-Robot Collaborative Assembly in Manufacturing Systems Using Large Language Models 2024/06 Collaboration Delegation
Enhancing the LLM-Based Robot Manipulation Through Human-Robot Collaboration 2024/06 Collaboration Delegation
WebCanvas: Benchmarking Web Agents in Online Environments 2024/06 Collaboration Delegation
REVECA: Adaptive Planning and Trajectory-based Validation in Cooperative Language Agents Using Information Relevance and Relative Proximity 2024/05 Collaboration Delegation
A Human-Computer Collaborative Tool for Training a Single Large Language Model Agent into a Network through Few Examples 2024/04 Collaboration Delegation, Supervision
AgentCoord: Visually Exploring Coordination Strategy for LLM-based Multi-Agent Collaboration 2024/04 GitHub stars Collaboration Coordination
An LLM-based approach for Enabling Seamless Human-Robot Collaboration in Assembly 2024/04 Collaboration Delegation
Autonomous Evaluation and Refinement of Digital Agents 2024/04 GitHub stars Collaboration Delegation
PDFChatAnnotator: A Human-LLM Collaborative Multi-Modal Data Annotation Tool for PDF-Format Catalogs 2024/04 Collaboration Delegation
Embodied LLM Agents Learn to Cooperate in Organized Teams 2024/03 GitHub stars Collaboration Delegation
Large Language Model-based Human-Agent Collaboration for Complex Task Solving 2024/02 GitHub stars Collaboration Delegation
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue 2024/02 GitHub stars Collaboration Delegation
A2C: A Modular Multi-stage Collaborative Decision Framework for Human-AI Teams 2024/01 Collaboration Coordination
Ask-before-Plan: Proactive Language Agents for Real-World Planning 2024/01 GitHub stars Collaboration Coordination
LLM-Powered Hierarchical Language Agent for Real-time Human-AI Coordination 2023/12 GitHub stars Collaboration Supervision, Delegation
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents 2023/10 Collaboration, Competition, Coopetition Coordination
Drive As You Speak: Enabling Human-Like Interaction With Large Language Models in Autonomous Vehicles 2023/09 Collaboration Delegation
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback 2023/09 Collaboration Delegation
MindAgent: Emergent Gaming Interaction 2023/09 GitHub stars Collaboration Coordination
LLM-Based Human-Robot Collaboration Framework for Manipulation Tasks 2023/08 Collaboration Supervision, Delegation
MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework 2023/08 GitHub stars Collaboration Coordination
Building Cooperative Embodied Agents Modularly with Large Language Models 2023/07 Collaboration Cooperation
Embodied Task Planning with Large Language Models 2023/07 GitHub stars Collaboration Delegation
Improved Trust in Human-Robot Collaboration With ChatGPT 2023/06 Collaboration Delegation
Investigating Agency of LLMs in Human-AI Collaboration Tasks 2023/05 Collaboration Cooperation, Delegation
Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents through Help Feedback 2023/04 Collaboration Delegation
PaLM-E: An Embodied Multimodal Language Model 2023/03 Collaboration Delegation
AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts 2021/10 Collaboration Delegation

🎛️ Orchestration

(©️click here back to table of contents👆🏻)

Title Date & Code Orchestration Strategy Orchestration Synchronization
Just A Rather Very Intelligent Spoken Agent (JarvisBench) 2026/07 Simultaneous Asynchronous
LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability 2026/07 GitHub stars One-by-One Synchronous
HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation 2026/07 One-by-One Synchronous
PerspectiveGap: A Benchmark for Multi-Agent Orchestration Prompting 2026/06 GitHub stars Simultaneous Asynchronous
Uncertainty Decomposition for Clarification Seeking in LLM Agents 2026/06 One-by-One Synchronous
Learning User Simulators with Turing Rewards 2026/06 One-by-One Synchronous
Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents (TRACE) 2026/06 GitHub stars One-by-One Synchronous
Re-Centering Humans in LLM Personalization 2026/06 One-by-One Asynchronous
CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive Engagement 2026/06 Simultaneous Synchronous
Human Oversight of Agentic Systems in Practice: Examining the Oversight Work, Challenges, and Heuristics of Developers Using Software Agents 2026/06 One-by-One Asynchronous
Uncertainty-Aware Clarification in LLM Agents with Information Gain 2026/06 One-by-One Synchronous
Not All Uncertainty Is Equal: How Uncertainty Granularity Shapes Human Verification in LLM-Assisted Decision Making 2026/05 One-by-One Synchronous
VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions 2026/05 GitHub stars One-by-One Synchronous
Reinforcing Human Behavior Simulation via Verbal Feedback (DITTO & SOUL) 2026/05 One-by-One Synchronous
ECHO: Explainable Co-editing with Human-in-the-loop Operations for Presentation Refinement 2026/05 One-by-One Synchronous
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators 2026/05 One-by-One Synchronous
SalesSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators 2026/05 One-by-One Synchronous
A Decoupled Human-in-the-Loop System for Controlled Autonomy in Agentic Workflows 2026/04 One-by-One Asynchronous
CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks 2026/04 One-by-One Synchronous
When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation (InterruptBench) 2026/04 GitHub stars One-by-One Synchronous
Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents 2026/03 One-by-One Synchronous
User Preference Modeling for Conversational LLM Agents: Weak Rewards from Retrieval-Augmented Interaction (VARS) 2026/03 GitHub stars One-by-One Synchronous
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks 2026/03 GitHub stars One-by-One Asynchronous
AgentDS: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science 2026/03 One-by-One Asynchronous
ViviDoc: Generating Interactive Documents through Human-Agent Collaboration 2026/03 GitHub stars One-by-One Synchronous
InfoPO: Information-Driven Policy Optimization for User-Centric Agents 2026/02 GitHub stars One-by-One Synchronous
Modeling Distinct Human Interaction in Web Agents (CowCorpus) 2026/02 Simultaneous Synchronous
Overseeing Agents Without Constant Oversight: Challenges and Opportunities 2026/02 One-by-One Asynchronous
Learning Personalized Agents from Human Feedback 2026/02 GitHub stars One-by-One Synchronous
Value of Information: A Framework for Human-Agent Communication 2026/01 One-by-One Synchronous
Progressive Ideation using an Agentic AI Framework for Human-AI Co-Creation (MIDAS) 2026/01 One-by-One Synchronous
From Correctness to Collaboration: Toward a Human-Centered Framework for Evaluating AI Agent Behavior in Software Engineering 2025/12 One-by-One Synchronous
CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding 2025/11 Simultaneous Synchronous
Training Proactive and Personalized LLM Agents 2025/11 GitHub stars One-by-One Synchronous
Training LLM Agents to Empower Humans 2025/10 GitHub stars One-by-One Synchronous
How can we assess human-agent interactions? Case studies in software agent design 2025/10 One-by-One Synchronous
RECODE-H: A Benchmark for Research Code Development with Interactive Human Feedback 2025/10 GitHub stars One-by-One Synchronous
UserRL: Training Proactive User-Centric Agent via Reinforcement Learning 2025/09 GitHub stars One-by-One Synchronous
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use 2025/08 GitHub stars One-by-One Synchronous
Magentic-UI: Towards Human-in-the-loop Agentic Systems 2025/07 GitHub stars Simultaneous Synchronous
UserBench: An Interactive Gym Environment for User-Centric Agents 2025/07 GitHub stars One-by-One Synchronous
Enabling Self-Improving Agents to Learn at Test Time With Human-In-The-Loop Guidance 2025/07 GitHub stars One-by-One Synchronous
τ2-Bench: Evaluating Conversational Agents in a Dual-Control Environment 2025/06 GitHub stars One-by-One Synchronous
Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training 2025/05 One-by-One Synchronous
Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild 2025/05 GitHub stars One-by-One Synchronous
XtraGPT: LLMs for Human-AI Collaboration on Controllable Academic Paper Revision 2025/05 GitHub stars One-by-One Synchronous
MineWorld: A Real-Time and Open-Source Interactive World Model on Minecraft 2025/04 GitHub stars One-by-One Synchronous
Experimental Exploration: Investigating Cooperative Interaction Behavior Between Humans and Large Language Model Agents 2025/03 One-by-One Synchronous
FinArena: A Human-Agent Collaboration Framework for Financial Market Analysis and Forecasting 2025/03 One-by-One Synchronous
SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks 2025/03 GitHub stars One-by-One Synchronous
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments 2025/02 GitHub stars One-by-One Asynchronous
Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI Collaboration 2025/02 GitHub stars Simultaneous Asynchronous
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation 2025/01 One-by-One Synchronous
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration (Co-Gym) 2024/12 GitHub stars One-by-One Asynchronous
Towards Modeling Human-Agentic Collaborative Workflows: A BPMN Extension 2024/12 GitHub stars One-by-One Asynchronous
To Help or Not to Help: LLM-based Attentive Support for Human-Robot Group Interactions 2024/10 GitHub stars One-by-One Synchronous
PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-Agent Tasks 2024/10 GitHub stars One-by-One Asynchronous
AssistantX: An LLM-Powered Proactive Assistant in Collaborative Human-Populated Environment 2024/09 GitHub stars One-by-One Asynchronous
Mutual Theory of Mind in Human-AI Collaboration: An Empirical Study with LLM-driven AI Agents in a Real-time Shared Workspace Task 2024/09 Simultaneous Synchronous
Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent Conversations 2024/08 GitHub stars One-by-One Synchronous
Human-LLM Collaboration in Generative Design for Customization 2024/07 One-by-One Synchronous
DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning 2024/06 GitHub stars One-by-One Asynchronous
Enhancing Human-Robot Collaborative Assembly in Manufacturing Systems Using Large Language Models 2024/06 One-by-One Synchronous
Enhancing the LLM-Based Robot Manipulation Through Human-Robot Collaboration 2024/06 One-by-One Synchronous
WebCanvas: Benchmarking Web Agents in Online Environments 2024/06 One-by-One Synchronous
REVECA: Adaptive Planning and Trajectory-based Validation in Cooperative Language Agents Using Information Relevance and Relative Proximity 2024/05 One-by-One Asynchronous
A Human-Computer Collaborative Tool for Training a Single Large Language Model Agent into a Network through Few Examples 2024/04 One-by-One Synchronous
AgentCoord: Visually Exploring Coordination Strategy for LLM-based Multi-Agent Collaboration 2024/04 GitHub stars One-by-One Synchronous
An LLM-based approach for Enabling Seamless Human-Robot Collaboration in Assembly 2024/04 One-by-One Synchronous
Autonomous Evaluation and Refinement of Digital Agents 2024/04 GitHub stars One-by-One Asynchronous
PDFChatAnnotator: A Human-LLM Collaborative Multi-Modal Data Annotation Tool for PDF-Format Catalogs 2024/04 One-by-One Synchronous
Embodied LLM Agents Learn to Cooperate in Organized Teams 2024/03 GitHub stars One-by-One Synchronous
Large Language Model-based Human-Agent Collaboration for Complex Task Solving 2024/02 GitHub stars One-by-One Synchronous
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue 2024/02 GitHub stars One-by-One Synchronous
A2C: A Modular Multi-stage Collaborative Decision Framework for Human-AI Teams 2024/01 One-by-One Asynchronous
Ask-before-Plan: Proactive Language Agents for Real-World Planning 2024/01 GitHub stars One-by-One Synchronous
LLM-Powered Hierarchical Language Agent for Real-time Human-AI Coordination 2023/12 GitHub stars Simultaneous Synchronous
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents 2023/10 One-by-One Synchronous
Drive As You Speak: Enabling Human-Like Interaction With Large Language Models in Autonomous Vehicles 2023/09 One-by-One Synchronous
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback 2023/09 One-by-One Synchronous
MindAgent: Emergent Gaming Interaction 2023/09 GitHub stars One-by-One Synchronous
LLM-Based Human-Robot Collaboration Framework for Manipulation Tasks 2023/08 One-by-One Synchronous
MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework 2023/08 GitHub stars One-by-One Asynchronous
Building Cooperative Embodied Agents Modularly with Large Language Models 2023/07 Simultaneous Synchronous
Embodied Task Planning with Large Language Models 2023/07 GitHub stars One-by-One Asynchronous
Improved Trust in Human-Robot Collaboration With ChatGPT 2023/06 One-by-One Synchronous
Investigating Agency of LLMs in Human-AI Collaboration Tasks 2023/05 One-by-One Synchronous
Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents through Help Feedback 2023/04 One-by-One Synchronous
PaLM-E: An Embodied Multimodal Language Model 2023/03 One-by-One Synchronous
AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts 2021/10 One-by-One Synchronous

💬 Communication

(©️click here back to table of contents👆🏻)

Title Date & Code Communication Structure Communication Mode
Just A Rather Very Intelligent Spoken Agent (JarvisBench) 2026/07 Hierarchical Conversation, Observation
LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability 2026/07 GitHub stars Decentralized Conversation
HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation 2026/07 Hierarchical, Centralized Conversation, Observation
Uncertainty Decomposition for Clarification Seeking in LLM Agents 2026/06 Centralized Conversation
Learning User Simulators with Turing Rewards 2026/06 Decentralized Conversation
Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents (TRACE) 2026/06 GitHub stars Hierarchical Conversation
Re-Centering Humans in LLM Personalization 2026/06 Centralized Conversation
CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive Engagement 2026/06 Decentralized Conversation, Observation
Human Oversight of Agentic Systems in Practice: Examining the Oversight Work, Challenges, and Heuristics of Developers Using Software Agents 2026/06 Hierarchical Observation, Conversation
Uncertainty-Aware Clarification in LLM Agents with Information Gain 2026/06 Centralized Conversation
Not All Uncertainty Is Equal: How Uncertainty Granularity Shapes Human Verification in LLM-Assisted Decision Making 2026/05 Centralized Conversation
VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions 2026/05 GitHub stars Centralized Conversation
Reinforcing Human Behavior Simulation via Verbal Feedback (DITTO & SOUL) 2026/05 Decentralized Conversation
ECHO: Explainable Co-editing with Human-in-the-loop Operations for Presentation Refinement 2026/05 Hierarchical Conversation, Observation
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators 2026/05 Decentralized Message Pool
SalesSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators 2026/05 Centralized Conversation, Observation
A Decoupled Human-in-the-Loop System for Controlled Autonomy in Agentic Workflows 2026/04 Hierarchical Conversation
CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks 2026/04 Centralized Conversation
When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation (InterruptBench) 2026/04 GitHub stars Centralized Conversation
Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents 2026/03 Hierarchical Conversation
User Preference Modeling for Conversational LLM Agents: Weak Rewards from Retrieval-Augmented Interaction (VARS) 2026/03 GitHub stars Centralized Conversation
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks 2026/03 GitHub stars Centralized Observation, Conversation
AgentDS: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science 2026/03 Centralized Conversation
ViviDoc: Generating Interactive Documents through Human-Agent Collaboration 2026/03 GitHub stars Hierarchical Conversation
InfoPO: Information-Driven Policy Optimization for User-Centric Agents 2026/02 GitHub stars Centralized Conversation
Modeling Distinct Human Interaction in Web Agents (CowCorpus) 2026/02 Centralized Observation, Conversation
Overseeing Agents Without Constant Oversight: Challenges and Opportunities 2026/02 Hierarchical Observation
Learning Personalized Agents from Human Feedback 2026/02 GitHub stars Centralized Conversation
Value of Information: A Framework for Human-Agent Communication 2026/01 Centralized Conversation
Progressive Ideation using an Agentic AI Framework for Human-AI Co-Creation (MIDAS) 2026/01 Decentralized Conversation
From Correctness to Collaboration: Toward a Human-Centered Framework for Evaluating AI Agent Behavior in Software Engineering 2025/12 Hierarchical Conversation
CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding 2025/11 Centralized Conversation, Observation
Training Proactive and Personalized LLM Agents 2025/11 GitHub stars Decentralized Conversation
Training LLM Agents to Empower Humans 2025/10 GitHub stars Hierarchical Conversation, Observation
How can we assess human-agent interactions? Case studies in software agent design 2025/10 Hierarchical Conversation
RECODE-H: A Benchmark for Research Code Development with Interactive Human Feedback 2025/10 GitHub stars Hierarchical Conversation
UserRL: Training Proactive User-Centric Agent via Reinforcement Learning 2025/09 GitHub stars Hierarchical Conversation
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use 2025/08 GitHub stars Hierarchical Conversation
Magentic-UI: Towards Human-in-the-loop Agentic Systems 2025/07 GitHub stars Hierarchical, Centralized Conversation, Observation
UserBench: An Interactive Gym Environment for User-Centric Agents 2025/07 GitHub stars Decentralized Conversation
Enabling Self-Improving Agents to Learn at Test Time With Human-In-The-Loop Guidance 2025/07 GitHub stars Hierarchical Conversation
τ2-Bench: Evaluating Conversational Agents in a Dual-Control Environment 2025/06 GitHub stars Decentralized Conversation, Observation
Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training 2025/05 Decentralized Conversation
Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild 2025/05 GitHub stars Centralized Conversation
XtraGPT: LLMs for Human-AI Collaboration on Controllable Academic Paper Revision 2025/05 GitHub stars Centralized Conversation
MineWorld: A Real-Time and Open-Source Interactive World Model on Minecraft 2025/04 GitHub stars Decentralized Conversation
Experimental Exploration: Investigating Cooperative Interaction Behavior Between Humans and Large Language Model Agents 2025/03 Decentralized Conversation
FinArena: A Human-Agent Collaboration Framework for Financial Market Analysis and Forecasting 2025/03 Hierarchical Conversation
SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks 2025/03 GitHub stars Decentralized Conversation
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments 2025/02 GitHub stars Decentralized Conversation
Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI Collaboration 2025/02 GitHub stars Decentralized Observation
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation 2025/01 Decentralized Conversation
Towards Modeling Human-Agentic Collaborative Workflows: A BPMN Extension 2024/12 GitHub stars Decentralized Conversation
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration (Co-Gym) 2024/12 GitHub stars Decentralized Conversation
To Help or Not to Help: LLM-based Attentive Support for Human-Robot Group Interactions 2024/10 GitHub stars Decentralized Observation
PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-Agent Tasks 2024/10 GitHub stars Decentralized Conversation
AssistantX: An LLM-Powered Proactive Assistant in Collaborative Human-Populated Environment 2024/09 GitHub stars Decentralized Message Pool
Mutual Theory of Mind in Human-AI Collaboration: An Empirical Study with LLM-driven AI Agents in a Real-time Shared Workspace Task 2024/09 Decentralized Conversation
Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent Conversations 2024/08 GitHub stars Decentralized Conversation
Human-LLM Collaboration in Generative Design for Customization 2024/07 Decentralized Conversation
DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning 2024/06 GitHub stars Decentralized Conversation
Enhancing Human-Robot Collaborative Assembly in Manufacturing Systems Using Large Language Models 2024/06 Decentralized Conversation
Enhancing the LLM-Based Robot Manipulation Through Human-Robot Collaboration 2024/06 Decentralized Conversation
WebCanvas: Benchmarking Web Agents in Online Environments 2024/06 Decentralized Conversation
REVECA: Adaptive Planning and Trajectory-based Validation in Cooperative Language Agents Using Information Relevance and Relative Proximity 2024/05 Hierarchical Observation
A Human-Computer Collaborative Tool for Training a Single Large Language Model Agent into a Network through Few Examples 2024/04 Decentralized Conversation
AgentCoord: Visually Exploring Coordination Strategy for LLM-based Multi-Agent Collaboration 2024/04 GitHub stars Decentralized Conversation
An LLM-based approach for Enabling Seamless Human-Robot Collaboration in Assembly 2024/04 Centralized Conversation
Autonomous Evaluation and Refinement of Digital Agents 2024/04 GitHub stars Decentralized Conversation
PDFChatAnnotator: A Human-LLM Collaborative Multi-Modal Data Annotation Tool for PDF-Format Catalogs 2024/04 Decentralized Conversation
Embodied LLM Agents Learn to Cooperate in Organized Teams 2024/03 GitHub stars Decentralized Conversation
Large Language Model-based Human-Agent Collaboration for Complex Task Solving 2024/02 GitHub stars Decentralized Conversation
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue 2024/02 GitHub stars Decentralized Conversation
A2C: A Modular Multi-stage Collaborative Decision Framework for Human-AI Teams 2024/01 Decentralized Conversation
Ask-before-Plan: Proactive Language Agents for Real-World Planning 2024/01 GitHub stars Centralized Conversation
LLM-Powered Hierarchical Language Agent for Real-time Human-AI Coordination 2023/12 GitHub stars Hierarchical Conversation
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents 2023/10 Centralized Conversation
Drive As You Speak: Enabling Human-Like Interaction With Large Language Models in Autonomous Vehicles 2023/09 Centralized Conversation
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback 2023/09 Decentralized Conversation
MindAgent: Emergent Gaming Interaction 2023/09 GitHub stars Centralized Conversation
LLM-Based Human-Robot Collaboration Framework for Manipulation Tasks 2023/08 Decentralized Conversation
MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework 2023/08 GitHub stars Decentralized Message Pool
Building Cooperative Embodied Agents Modularly with Large Language Models 2023/07 Decentralized Conversation
Embodied Task Planning with Large Language Models 2023/07 GitHub stars Decentralized Conversation
Improved Trust in Human-Robot Collaboration With ChatGPT 2023/06 Decentralized Conversation
Investigating Agency of LLMs in Human-AI Collaboration Tasks 2023/05 Decentralized Conversation
Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents through Help Feedback 2023/04 Decentralized Conversation
PaLM-E: An Embodied Multimodal Language Model 2023/03 Decentralized Conversation
AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts 2021/10 Hierarchical Conversation

📌 Contributing

(©️click here back to table of contents👆🏻)

Contributions are welcome! If you have relevant papers, code, or insights, feel free to submit a request 🤗.

📝 Citation

(©️click here back to table of contents👆🏻)

If you find this repository useful, please consider citing our papers 💕:

The survey has been accepted to ACL 2026; the BibTeX below will be updated to the proceedings entry once it is available.

@misc{zou2025llmbasedhumanagentcollaborationinteraction,
      title={LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey}, 
      author={Henry Peng Zou and Wei-Chieh Huang and Yaozu Wu and Yankai Chen and Chunyu Miao and Hoang Nguyen and Yue Zhou and Weizhi Zhang and Liancheng Fang and Langzhou He and Yangning Li and Dongyuan Li and Renhe Jiang and Xue Liu and Philip S. Yu},
      year={2025},
      eprint={2505.00753},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2505.00753}, 
}

@misc{zou2025collaborativeintelligencehumanagentsystems,
      title={A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy}, 
      author={Henry Peng Zou and Wei-Chieh Huang and Yaozu Wu and Chunyu Miao and Dongyuan Li and Aiwei Liu and Yue Zhou and Yankai Chen and Weizhi Zhang and Yangning Li and Liancheng Fang and Renhe Jiang and Philip S. Yu},
      year={2025},
      eprint={2506.09420},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2506.09420}, 
}

⭐ Star History

Star History Chart

About

[ACL 2026] LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey | Awesome Human-Agent Collaboration | Human-AI Collaboration

Topics

Resources

Stars

234 stars

Watchers

9 watching

Forks

Releases

Packages

Used by

Contributors