Journal Article10.48550/arXiv.2212.09746
Evaluating Human-Language Model Interaction
Mina Lee,Megha Srivastava,Amelia Hardy,John Thickstun,Esin Durmus,Ashwin Paranjape,Ines Gerard-Ursin,Xing Li,Faisal Ladhak,Frieda Rong,Rose E. Wang,Minae Kwon,Joon Sung Park,Hancheng Cao,Tony Lee,Rishi Bommasani,Michael S. Bernstein,Percy Liang +17 more
TL;DR: The authors developed a new framework, Human-AI Language-based Interaction Evaluation (HALIE), that defines the components of interactive systems and dimensions to consider when designing evaluation metrics.
read more
Abstract: Many real-world applications of language models (LMs), such as writing assistance and code autocomplete, involve human-LM interaction. However, most benchmarks are non-interactive in that a model produces output without human involvement. To evaluate human-LM interaction, we develop a new framework, Human-AI Language-based Interaction Evaluation (HALIE), that defines the components of interactive systems and dimensions to consider when designing evaluation metrics. Compared to standard, non-interactive evaluation, HALIE captures (i) the interactive process, not only the final output; (ii) the first-person subjective experience, not just a third-party assessment; and (iii) notions of preference beyond quality (e.g., enjoyment and ownership). We then design five tasks to cover different forms of interaction: social dialogue, question answering, crossword puzzles, summarization, and metaphor generation. With four state-of-the-art LMs (three variants of OpenAI's GPT-3 and AI21 Labs' Jurassic-1), we find that better non-interactive performance does not always translate to better human-LM interaction. In particular, we highlight three cases where the results from non-interactive and interactive metrics diverge and underscore the importance of human-LM interaction for LM evaluation.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Holistic Evaluation of Language Models
Percy Liang,Rishi Bommasani,Tony Lee,Dimitris Tsipras,Dilara Soylu,Michihiro Yasunaga,Yian Zhang,Deepak Narayanan,Yuhuai Wu,Ananya Kumar,Benjamin Newman,Binhang Yuan,Bobby Yan,Ce Zhang,Christian Cosgrove,Christopher D. Manning,Christopher R'e,Diana Acosta-Navas,Drew A. Hudson,Eric Zelikman,Esin Durmus,Faisal Ladhak,Frieda Rong,Hongyu Ren,Huaxiu Yao,Jue Wang,Keshav Santhanam,Laurel Orr,Lucia Zheng,Byron Rogers,Mirac Suzgun,Nathan S. Kim,Neel Guha,Niladri S. Chatterji,Peter Henderson,Qian Huang,Ryan Chi,Michael Xie,Shibani Santurkar,Surya Ganguli,Tatsunori Hashimoto,Thomas Icard,Tianyi Zhang,Vishrav Chaudhary,William Wang,Xuechen Li,Yifan Mai,Yuhui Zhang,Yuta Koreeda +48 more
TL;DR: The Holistic Evaluation of Language Models (HELM) as mentioned in this paper ) is a popular benchmark for language models, with 30 models evaluated on 16 core scenarios and 7 metrics, exposing important trade-offs.
582
A Survey on Large Language Model based Autonomous Agents
Lei Wang,Cheng-jian Ma,Xueyang Feng,Zeyu Zhang,Hao-ran Yang,Jingsen Zhang,Zhi-Yang Chen,Jiakai Tang,Xu Chen,Yankai Lin,Wayne Xin Zhao,Zhewei Wei,Ji-Rong Wen +12 more
TL;DR: A systematic review of the field of LLM-based autonomous agents from a holistic perspective, and proposes a unified framework that encompasses a majority of the previous work.
BloombergGPT: A Large Language Model for Finance
Shijie Wu,Ozan Irsoy,Steven Lu,Mark Dredze,Sebastian Gehrmann,Prabhanjan Kambadur,D Rosenberg,Gideon Mann +7 more
TL;DR: The authors presented a 50 billion parameter language model that is trained on a wide range of financial data, including a 363 billion token dataset, augmented with 345 billion tokens from general purpose datasets.
425
Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
Kai Greshake,Sahar Abdelnabi,Shailesh Nandan Mishra,Christoph Endres,T Holz,Mario Fritz +5 more
- 23 Feb 2023
TL;DR: In this article , the authors reveal new attack vectors, using indirect prompt injection, that enable adversaries to remotely (without a direct interface) exploit LLM-integrated applications by strategically injecting prompts into data likely to be retrieved.
Peer Review
Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction
Renee Shelby,Shalaleh Rismani,Kathryn Henne,AJung Moon,Negar Rostamzadeh,Paul Nicholas,N. Yilla,Jess Gallegos,Andrew Smart,Emilio Garcia,Gurleen Virk +10 more
- 11 Oct 2022
TL;DR: In this paper , the authors present an applied taxonomy of sociotechnical harms to support a more systematic surfacing of potential harms in algorithmic systems, including representational, allocative, quality-of-service, interpersonal harms, and social system/societal harms.
References
•Proceedings Article
Search in the Lost Sense of ``Query'': Question Formulation in Web Search Queries and its Temporal Changes
Bo Pang,Ravi Kumar +1 more
- 19 Jun 2011
TL;DR: Through a systematic, large-scale study, it is found that as time goes by, web users are more likely to use questions to express their search intent.
23
Dr.Fill: Crosswords and an Implemented Solver for Singly Weighted CSPs
TL;DR: Dr. Fill is described, a program that solves American-style crossword puzzles by converting crosswords to weighted csps, and then using a variety of novel techniques to find a solution.
16
Alexa, Let's Work Together: Introducing the First Alexa Prize TaskBot Challenge on Conversational Task Assistance
Anna Gottardi,Osman Ipek,Giuseppe Castellucci,Shui Hu,Lavina Vaz,Yao Lu,Anju Khatri,Anjali Chadha,Desheng Zhang,Sattvik Sahai,Prerna Dwivedi,Hangjie Shi,Lu Hu,Andy Huang,Lu Dai,Bo Yang,Varun Somani,Pankaj Rajan,Ron Rezac,Michael Johnston,Savanna Stiff,Leslie Ball,David Carmel,Yang Liu,Dilek Hakkani-Tur,Oleg Rokhlenko,Kate Bland,E. Agichtein,Reza Ghanadan,Yoelle Maarek +29 more
TL;DR: An overview of the TaskBot challenge is provided, the infrastructure support provided to the teams with the CoBot Toolkit is described, and the approaches the participating teams took to overcome the research challenges are summarized.
13
Interactive Query-Assisted Summarization via Deep Reinforcement Learning
Ori Shapira,Ramakanth Pasunuru,Mohit Bansal,Ido D. Dagan,Yael Amsterdamer +4 more
- 01 Jan 2022
TL;DR: Two novel deep reinforcement learning models are proposed for the interactive summarization task that address, respectively, the subtask of summarizing salient information that adheres to user queries, and the subtasking of listing suggested queries to assist users throughout their exploration.
9
Aligning Offline Metrics and Human Judgments of Value of AI-Pair Programmers
TL;DR: In this article , a user study with (N = 49 ) experienced programmers showed that while both correctness and effort correlate with value, the association is strongest for effort, and they argue that effort should be considered as an important dimension of evaluation in code generation scenarios.
9