Journal Article10.1145/3611643.3616283
Statistical Type Inference for Incomplete Programs
Yaohui Peng,Jing Xie,Qiongling Yang,Hanwen Guo,Qingan Li,Jingling Xue,Mengting Yuan +6 more
- 30 Nov 2023
1
TL;DR: Stir is a novel two-stage approach for inferring types in incomplete programs that may be ill-formed, where whole-program syntactic analysis often fails, and reduces it to a sequence-to-graph parsing problem.
read more
Abstract: We propose a novel two-stage approach, Stir, for inferring types in incomplete programs that may be ill-formed, where whole-program syntactic analysis often fails. In the first stage, Stir predicts a type tag for each token by using neural networks, and consequently, infers all the simple types in the program. In the second stage, Stir refines the complex types for the tokens with predicted complex type tags. Unlike existing machine-learning-based approaches, which solve type inference as a classification problem, Stir reduces it to a sequence-to-graph parsing problem. According to our experimental results, Stir achieves an accuracy of 97.37 % for simple types. By representing complex types as directed graphs (type graphs), Stir achieves a type similarity score of 77.36 % and 59.61 % for complex types and zero-shot complex types, respectively.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Inferring Pluggable Types with Machine Learning
Khan Siddiqui,Martin Kellogg +1 more
- 21 Jun 2024
TL;DR: Inferring pluggable types with machine learning is feasible and can significantly reduce the burden of writing type annotations in legacy codebases.
References
Using Pre-Trained Models to Boost Code Review Automation
Rosalia Tufano,Simone Masiero,Antonio Mastropaolo,Luca Pascarella,Denys Poshyvanyk,Gabriele Bavota +5 more
- 18 Jan 2022
TL;DR: It is demonstrated that a pre-trained Text-To-Text Transfer Transformer model can outperform previous DL models for automating code review tasks and is also conducted on a larger and more realistic dataset of code review activities.
Typilus: neural type hints
Miltiadis Allamanis,Earl T. Barr,Soline Ducousso,Zheng Gao +3 more
- 11 Jun 2020
TL;DR: Typilus as mentioned in this paper uses deep similarity learning to learn a continuous relaxation of the discrete space of types and embeds the type properties of a symbol (i.e. identifier) into it.
CodeFill: Multi-token Code Completion by Jointly learning from Structure and Naming Sequences
Maliheh Izadi,Roberta Gismondi,Georgios Gousios +2 more
- 14 Feb 2022
TL;DR: This work presents CodeFill, a language model for autocompletion that combines learned structure and naming information and is trained both for single-token and multi-token prediction, which enables it to learn long-range dependencies among grammatical and naming elements.
•Posted Content
TypeWriter: Neural Type Prediction with Search-based Validation
TL;DR: TypeWriter is presented, the first combination of probabilistic type prediction with search-based refinement of predicted types, which can fully annotate between 14% to 44% of the files in a randomly selected corpus, while ensuring type correctness.
TypeWriter: neural type prediction with search-based validation
Michael Pradel,Georgios Gousios,Jason Liu,Satish Chandra +3 more
- 08 Nov 2020
TL;DR: TypeWriter as mentioned in this paper combines probabilistic type prediction with search-based refinement of predicted types to annotate large code bases written in dynamically typed languages, such as JavaScript or Python, by combining natural language properties of code with programming language-level information.
80