From BERT to GPT-3 codex
TL;DR: The goal of the tutorial is to introduce database researchers to the latest generation of language models, and to their use cases in the domain of data management.
read more
Abstract: Large language models have recently advanced the state of the art on many natural language processing benchmarks. The newest generation of models can be applied to a variety of tasks with little to no specialized training. This technology creates various opportunities for applications in the context of data management.
The tutorial will introduce participants to basic background on language models, discuss different methods to use language models, and give an overview and short demonstration of available libraries and APIs. Models for generating natural language will be considered as well as models, such as GPT-3 Codex, which complete program code or generate code from natural language instructions. Finally, the tutorial will discuss recent research in the database community that exploits language models in the context of traditional database systems or proposes novel system architectures that are based on them.
The tutorial is targeted at database researchers. No prior background on language models is required. The goal of the tutorial is to introduce database researchers to the latest generation of language models, and to their use cases in the domain of data management.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
ChatGPT: Fundamentals, Applications and Social Impacts
29 Nov 2022
TL;DR: ChatGPT as mentioned in this paper is a natural language processing model developed in 2022 by OpenAI for open-ended conversations, which is based on GPT-3.5, the third-generation model from OpenAI and can power conversational AI applications like virtual assistants and chatbots.
166
A Universal Question-Answering Platform for Knowledge Graphs
Reham Omar,Ishika Dhall,Panos Kalnis,Essam Mansour +3 more
- 01 Mar 2023
TL;DR: KGQAn as mentioned in this paper is a universal QA system that does not need to be tailored to each target knowledge graph and uses a neural sequence-to-sequence model to convert a question into an intermediate abstract representation.
Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity
Haojun Xia,Zheng Zhang,Yuchao Li,Donglin Zhuang,Zhongzhu Zhou,Xiang Qiu,Yong Li,Wei Lin,Shuaiwen Leon Song +8 more
TL;DR: Flash-LLM enables cost-effective and highly-efficient large generative model inference with unstructured sparsity on tensor cores, significantly reducing GPU memory consumption and computation cost while maintaining accuracy.
9
Can Large Language Models Predict Data Correlations from Column Names?
TL;DR: The paper introduces a novel benchmark for data correlation analysis, created by analyzing thousands of Kaggle data sets, and uses that data to study the ability of language models to predict correlation, based on column names.
8
References
•Proceedings Article
Attention is All you Need
Ashish Vaswani,Noam Shazeer,Niki Parmar,Jakob Uszkoreit,Llion Jones,Aidan N. Gomez,Lukasz Kaiser,Illia Polosukhin +7 more
- 12 Jun 2017
TL;DR: This paper proposed a simple network architecture based solely on an attention mechanism, dispensing with recurrence and convolutions entirely and achieved state-of-the-art performance on English-to-French translation.
•Posted Content
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
TL;DR: A new language representation model, BERT, designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers, which can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks.
81.7K
•Posted Content
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu,Myle Ott,Naman Goyal,Jingfei Du,Mandar Joshi,Danqi Chen,Omer Levy,Michael Lewis,Luke Zettlemoyer,Veselin Stoyanov +9 more
TL;DR: It is found that BERT was significantly undertrained, and can match or exceed the performance of every model published after it, and the best model achieves state-of-the-art results on GLUE, RACE and SQuAD.
•Proceedings Article
Language Models are Few-Shot Learners
Tom B. Brown,Benjamin Mann,Nick Ryder,Melanie Subbiah,Jared Kaplan,Prafulla Dhariwal,Arvind Neelakantan,Pranav Shyam,Girish Sastry,Amanda Askell,Sandhini Agarwal,Ariel Herbert-Voss,Gretchen Krueger,Thomas Henighan,Rewon Child,Aditya Ramesh,Daniel M. Ziegler,Jeffrey Wu,Clemens Winter,Christopher Hesse,Mark Chen,Eric Sigler,Mateusz Litwin,Scott Gray,Benjamin Chess,Jack Clark,Christopher Berner,Samuel McCandlish,Alec Radford,Ilya Sutskever,Dario Amodei +30 more
- 28 May 2020
TL;DR: GPT-3 achieves strong performance on many NLP datasets, including translation, question-answering, and cloze tasks, as well as several tasks that require on-the-fly reasoning or domain adaptation, such as unscrambling words, using a novel word in a sentence, or performing 3-digit arithmetic.
•Posted Content
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel,Noam Shazeer,Adam Roberts,Katherine Lee,Sharan Narang,Michael Matena,Yanqi Zhou,Wei Li,Peter J. Liu +8 more
TL;DR: This systematic study compares pre-training objectives, architectures, unlabeled datasets, transfer approaches, and other factors on dozens of language understanding tasks and achieves state-of-the-art results on many benchmarks covering summarization, question answering, text classification, and more.
Related Papers (5)
Aarne Ranta
- 20 Aug 2014
S. Gusev,Andrey Chepovskiy +1 more
- 01 Jan 2011
Cristina Barros,Elena Lloret +1 more
- 01 Sep 2017