Reference. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Language models demonstrate both quantitative improvement and new qualitative capabilities with increasing scale. Despite their potentially transformative impact, these new capabilities are as yet poorly characterized. In order to inform future research, prepare for disruptive new model capabilities, and ameliorate socially harmful effects, it is vital that we understand the present and near-future capabilities and limitations of language models. To address this challenge, we introduce the Beyond the Imitation Game benchmark (BIG-bench). BIG-bench currently consists of 204 tasks, contributed by 450 authors across 132 institutions. Task topics are diverse, drawing problems from linguistics, childhood development, math, common-sense reasoning, biology, physics, social bias, software development, and beyond. BIG-bench focuses on tasks that are believed to be beyond the capabilities of current language models. We evaluate the behavior of OpenAI’s GPT models, Google-internal dense transformer architectures, and Switch-style sparse transformers on BIG-bench, across model sizes spanning millions to hundreds of billions of parameters. In addition, a team of human expert raters performed all tasks in order to provide a strong baseline. Findings include: model performance and calibration both improve with scale, but are poor in absolute terms (and when compared with rater performance); performance is remarkably similar across model classes, though with benefits from sparsity; tasks that improve gradually and predictably commonly involve a large knowledge or memorization component, whereas tasks that exhibit “breakthrough” behavior at a critical scale often involve multiple steps or components, or brittle metrics; social bias typically increases with scale in settings with ambiguous context, but this can be improved with prompting.
Cite
Cites 696 works (0 here)
External (696)
- Neural language models are effective plagiarists (2022)
- GPT-NeoX-20B: An open-source autoregressive language model (2022)
- PaLM: Scaling language modeling with pathways (2022)
- Unified scaling laws for routed language models (2022)
- Predictability and surprise in large generative models (2022)
- EleutherAI/lm-evaluation-harness: v0.2.0, March 2022 (2022)
- Training compute-optimal large language models (2022)
- Quality at a glance: An audit of web-crawled multilingual datasets (2022)
- Documenting geographically and contextually diverse data sources: The BigScience catalogue of language data and resources (2022)
- Multitask prompted training enables zero-shot task generalization (2022)
- Zero-shot recommendation as language modeling (2022)
- You reap what you sow: On the challenges of bias evaluation under multilingual settings (2022)
- LaMDA: Language models for dialog applications (2022)
- Chain of thought prompting elicits reasoning in large language models (2022)
- Designing effective sparse expert models (2022)
- Persistent anti-Muslim bias in large language models (2021)
- Efficient large scale language modeling with mixtures of experts (2021)
- A general language assistant as a laboratory for alignment (2021)
- Program synthesis with large language models (2021)
- Explaining neural scaling laws (2021)
- On the dangers of stochastic parrots: Can language models be too big? (2021)
- Think you have solved direct-answer question answering? Try ARC-DA, the direct-answer AI2 reasoning challenge (2021)
- Multimodal datasets: Misogyny, pornography, and malignant stereotypes (2021)
- On the opportunities and risks of foundation models (2021)
- What will it take to fix benchmarking in natural language understanding? (2021)
- Extracting training data from large language models (2021)
- Evaluating large language models trained on code (2021)
- Training verifiers to solve math word problems (2021)
- White Chicago cops use force more often than Black officers (2021)
- NL-Augmenter: A framework for task-sensitive natural language augmentation (2021)
- GLaM: Efficient scaling of language models with mixture-of-experts (2021)
- Neural path hunter: Reducing hallucination in dialogue systems via path grounding (2021)
- Cryptonite: A cryptic crossword benchmark for extreme ambiguity in language (2021)
- Measuring and improving consistency in pretrained language models (2021)
- Beyond English-centric multilingual machine translation (2021)
- Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity (2021)
- The pile: An 800GB dataset of diverse text for language modeling (2021)
- The GEM benchmark: Natural language generation, its evaluation and metrics (2021)
- Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies (2021)
- ePiC: Employing proverbs in context as a benchmark for abstract language understanding (2021)
- Disfl-QA: A benchmark dataset for understanding disfluencies in question answering (2021)
- A survey on recent approaches for natural language processing in low-resource scenarios (2021)
- Measuring coding challenge competence with APPS, 2021a (2021)
- Measuring massive multitask language understanding (2021)
- Measuring mathematical problem solving with the MATH dataset, 2021c (2021)
- Scaling laws for transfer (2021)
- National name report 2020 (2021)
- Alignment of language agents (2021)
- Dynabench: Rethinking benchmarking in NLP (2021)
- MultiEmo: Multilingual, multilevel, multidomain sentiment analysis corpus of consumer reviews (2021)
- Counterlogicals as counterconventionals (2021)
- Hurdles to progress in long-form question answering (2021)
- Can RNNs learn recursive nested subject-verb agreements?, 2021a (2021)
- Mechanisms for handling nested dependencies in neural-network language models and humans (2021)
- Towards few-shot fact-checking via perplexity (2021)
- Investigating memorization of conspiracy theories in text generation (2021)
- Riddlesense: Reasoning about riddle questions featuring linguistic creativity and commonsense knowledge, 2021a (2021)
- TruthfulQA: Measuring how models mimic human falsehoods, 2021b (2021)
- What makes good in-context examples for GPT-3?, 2021a (2021)
- Can small and synthetic benchmarks drive modeling innovation? A retrospective study of question answering modeling approaches, 2021b (2021)
- A token-level reference-free hallucination detection benchmark for free-form text generation, 2021c (2021)
- Unicorn on rainbow: A universal commonsense reasoning model on a new multitask benchmark (2021)
- What's in the box? An analysis of undesirable content in the Common Crawl corpus (2021)
- EventPlus: A temporal event understanding pipeline (2021)
- Few-shot bot: Prompt-based learning for dialogue systems (2021)
- Inclusive data visualization for people with disabilities: A call to action (2021)
- Research community dynamics behind popular AI benchmarks (2021)
- Recipe1M+: A dataset for learning cross-modal embeddings for cooking recipes and food images (2021)
- Acquisition of chess knowledge in AlphaZero (2021)
- Cross-task generalization via natural language crowdsourcing instructions (2021)
- Structure here, bias there: Hierarchical generalization by jointly learning syntactic transformations (2021)
- Deep double descent: Where bigger models and more data hurt (2021)
- Show your work: Scratchpads for intermediate computation with language models (2021)
- Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics (2021)
- A review of speaker diarization: Recent advances with deep learning (2021)
- BBQ: A hand-built bias benchmark for question answering (2021)
- Are NLP models really able to solve simple math word problems? (2021)
- Carbon emissions and large neural network training (2021)
- True few-shot learning with language models (2021)
- Few-shot instruction prompts for pretrained language models to detect social biases (2021)
- TIMEDIAL: Temporal commonsense reasoning in dialog (2021)
- Scaling language models: Methods, analysis & insights from training Gopher (2021)
- Med-BERT: Pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction (2021)
- Decrypting cryptic crosswords: Semantically complex wordplay puzzles as a target for NLP (2021)
- XTREME-R: Towards more challenging and nuanced multilingual evaluation (2021)
- Symbolic behaviour in artificial intelligence (2021)
- Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in NLP (2021)
- Get your vitamin C! Robust fact verification with contrastive evidence (2021)
- Towards causal representation learning (2021)
- Does he wink or does he nod? A challenging benchmark for evaluating word understanding of language models (2021)
- Counterfactual learning in networks: An empirical study of model dependence (2021)
- Retrieval augmentation reduces hallucination in conversation (2021)
- COM2SENSE: A commonsense reasoning benchmark with complementary sentences (2021)
- Early detection of freeze damage in navel orange fruit using nondestructive low intensity ultrasound coupled with machine learning (2021)
- Watching a language model learning chess (2021)
- ChePT – applying deep neural transformer models to chess move prediction and self-commentary (2021)
- Understanding the capabilities, limitations, and societal impact of large language models (2021)
- Representing numbers in NLP: A survey and a vision (2021)
- Metaphor paraphrasing and word sense disambiguation: Toward a new approach to automated metaphor (2021)
- Recent advances in neural metaphor processing: A linguistic, cognitive and social perspective (2021)
- Learning chess blindfolded: Evaluating language models on state tracking (2021)
- GPT-J-6B: A 6 billion parameter autoregressive language model, May 2021 (2021)
- Finetuned language models are zero-shot learners (2021)
- Language models are few-shot multilingual learners (2021)
- An embedding method for unseen words considering contextual information and morphological information (2021)
- The causal-neural connection: Expressiveness, learnability, and inference (2021)
- On hallucination and predictive uncertainty in conditional language generation (2021)
- ByT5: Towards a token-free future with pre-trained byte-to-byte models, 2021a (2021)
- mT5: A massively multilingual pre-trained text-to-text transformer (2021)
- Calibrate before use: Improving few-shot performance of language models (2021)
- A survey of neural networks and formal languages (2020)
- Learning convex optimization models (2020)
- A very unlikely chess game (2020)
- Structural language models of code (2020)
- A survey on approaches to computational humor generation (2020)
- Bringing stories alive: Generating interactive fiction worlds (2020)
- ColBERT: Using BERT sentence embedding for humor detection (2020)
- On the cross-lingual transferability of monolingual representations (2020)
- Generating fact checking explanations (2020)
- The Pushshift Reddit dataset (2020)
- The relationship between inference skills and reading comprehension (2020)
- Climbing towards NLU: On meaning, form, and understanding in the age of data (2020)
- Critical thinking for language models (2020)
- On the ability and limitations of transformers to recognize formal languages (2020)
- On the practical ability of recurrent neural networks to recognize hierarchical languages (2020)
- The importance of suppressing domain style in authorship analysis (2020)
- PIQA: reasoning about physical commonsense in natural language (2020)
- Language (technology) is power: A critical survey of "bias" in NLP (2020)
- GPT-3 creative fiction (2020)
- Language models are few-shot learners (2020)
- Developing self-awareness in robots via inner speech (2020)
- Generative pretraining from pixels (2020)
- Transformers play chess (2020)
- Abstraction and reasoning challenge (2020)
- Transformers as soft reasoners over language (2020)
- Automated data transformation with inductive programming and dynamic background knowledge (2020)
- Learning higher-order logic programs (2020)
- Playing text-based games with common sense (2020)
- When redundancy is useful: A Bayesian approach to “overinformative” referring expressions (2020)
- Calibration of pre-trained transformers (2020)
- On measuring and mitigating biased inferences of word embeddings (2020)
- Learning syllogism with Euler neural-networks (2020)
- RoFT: A tool for evaluating human detection of machine-generated text (2020)
- To test machine comprehension, start by defining comprehension (2020)
- FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization (2020)
- How can self-attention networks recognize Dyck-n languages? (2020)
- Dreamcoder: Growing generalizable, interpretable knowledge with wake-sleep Bayesian program learning (2020)
- Text editing by command (2020)
- Humor detection via an internal and external neural network (2020)
- Go figure: A meta evaluation of factuality in summarization (2020)
- Neurosymbolic AI: The 3rd wave (2020)
- Evaluating models' local decision boundaries via contrast sets (2020)
- SyntaxGym: An online platform for targeted evaluation of language models (2020)
- RealToxicityPrompts: Evaluating neural toxic degeneration in language models (2020)
- Conversational implicatures in English dialogue: Annotated dataset (2020)
- Injecting numerical reasoning skills into language models (2020)
- Transformer feed-forward layers are key-value memories, 2020b (2020)
- Irony detection in a multilingual context (2020)
- A report on the 2020 sarcasm detection shared task (2020)
- Are neural open-domain dialog systems robust to speech recognition errors in the dialog history? An empirical study (2020)
- Theoretical limitations of self-attention in neural sequence models (2020)
- ECONET: Effective continual pretraining of language models for event temporal reasoning (2020)
- Policy-driven neural response generation for knowledge-grounded dialog systems (2020)
- Aligning AI with shared human values (2020)
- Scaling laws for autoregressive generative modeling (2020)
- 3D-DEEP: 3-dimensional deep-learning based on elevation patterns for road scene interpretation (2020)
- TaPas: Weakly supervised table parsing via pre-training (2020)
- RNNs can generate bounded hierarchical languages with optimal memory (2020)
- Bridging anaphora resolution as question answering (2020)
- National name report 2019 (2020)
- Compositionality decomposed: How do neural networks generalise? (2020)
- Automatic detection of generated text is easiest when humans are fooled (2020)
- Leveraging passage retrieval with generative models for open domain question answering (2020)
- Indic-transformers: An analysis of transformer language models for indian languages (2020)
- Learning to execute instructions in a Minecraft dialogue (2020)
- Are natural language inference models IMPPRESsive? Learning IMPlicature and PRESupposition (2020)
- How can we know what language models know? (2020)
- Robust encodings: A framework for combating adversarial typos (2020)
- Template guided text generation for task-oriented dialogue (2020)
- Scaling laws for neural language models (2020)
- Are pretrained language models symbolic reasoners over knowledge? (2020)
- ParsiNLU: A suite of language understanding challenges for persian, 2020a (2020)
- UNIFIEDQA: Crossing format boundaries with a single QA system (2020)
- Evaluating approaches to personalizing language models (2020)
- Against conventional wisdom (2020)
- All the news that’s fit to fabricate: AI-generated text as a tool of media misinformation (2020)
- Evaluating the factual consistency of abstractive text summarization (2020)
- Human vs. supervised machine learning: Who learns patterns faster? (2020)
- The NetHack learning environment (2020)
- Giving GPT-3 a Turing test (2020)
- Word meaning in minds and machines (2020)
- Language models as fact checkers? (2020)
- MLQA: Evaluating cross-lingual extractive question answering (2020)
- Retrieval-augmented generation for knowledge-intensive NLP tasks, 2020b (2020)
- Question and answer test-train overlap in open-domain question answering datasets, 2020c (2020)
- UNQOVERing stereotyping biases via underspecified questions (2020)
- DELPHI: Accurate deep ensemble model for protein interaction sites prediction (2020)
- Towards debiasing sentence representations (2020)
- Learning to contrast the counterfactual samples for robust visual question answering (2020)
- Birds have four legs?! NumerSense: probing numerical commonsense knowledge of pre-trained language models (2020)
- LogiQA: A challenge dataset for machine reading comprehension with logical reasoning, 2020a (2020)
- Interpretable multi-step reasoning with knowledge extraction on complex healthcare question answering, 2020b (2020)
- Multilingual denoising pre-training for neural machine translation, 2020c (2020)
- Scruples: A corpus of community ethical judgments on 32,000 real-life anecdotes (2020)
- Gender bias in neural natural language processing (2020)
- Language models as few-shot learner for task-oriented dialogue systems (2020)
- Low-resource languages: A review of past work and future challenges (2020)
- The next decade in AI: Four steps towards robust artificial intelligence (2020)
- GPT-3, bloviator: OpenAI’s language generator has no idea what it’s talking about (2020)
- OpenAI API alchemy: Emoji storytelling (2020)
- On faithfulness and factuality in abstractive summarization (2020)
- Does Syntax Need to Grow on Trees? Sources of Hierarchical Inductive Bias in Sequence-to-Sequence Networks (2020)
- USR: An unsupervised and reference free evaluation metric for dialog generation (2020)
- A framework for the computational linguistic analysis of dehumanization (2020)
- On the linguistic capacity of real-time counter automata (2020)
- The effect of natural distribution shift on question answering models (2020)
- StereoSet: Measuring stereotypical bias in pretrained language models (2020)
- The deep bootstrap framework: Good online learners are good offline generalizers (2020)
- Participatory research for low-resourced machine translation: A case study in African languages (2020)
- The chess transformer: Mastering play using generative language models (2020)
- iSarcasm: A dataset of intended sarcasm (2020)
- Sarcasm detection using context separators in online discourse (2020)
- Don't patronize me! An annotated dataset with patronizing and condescending language towards vulnerable communities (2020)
- Data cleaning: A case study with OpenRefine and Trifacta Wrangler (2020)
- Out of order: How important is the sequential order of words in a sentence in natural language understanding tasks? (2020)
- Generative language modeling for automated theorem proving (2020)
- A transformer-based approach to irony and sarcasm detection (2020)
- An analysis of the adaptation speed of causal models (2020)
- A survey on computational metaphor processing (2020)
- Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset (2020)
- Comparing conventions (2020)
- The child as hacker: Building more human-like models of learning (2020)
- The child as hacker (2020)
- How good is your tokenizer? On the monolingual performance of multilingual language models (2020)
- PuzzLing Machines: A challenge on learning from small data (2020)
- WINOGRANDE: An adversarial Winograd schema challenge at scale (2020)
- Masked language model scoring (2020)
- Social bias frames: Reasoning about social and power implications of language (2020)
- BLEURT: Learning robust metrics for text generation (2020)
- All your questions answered (2020)
- Neural logic reasoning (2020)
- EmoTag1200: Understanding the association between emojis and emotions (2020)
- DiscSense: Automated semantic analysis of discourse markers (2020)
- Learning to summarize from human feedback (2020)
- Metaphoric paraphrase generation (2020)
- Evolution and impact of bias in human and machine learning algorithm interaction (2020)
- Learning what makes a difference from counterfactual examples and gradient supervision (2020)
- Correlation-based network analysis combined with machine learning techniques highlight the role of the gaba shunt in brachypodium sylvaticum freezing tolerance (2020)
- Temporal reasoning in natural language inference (2020)
- Fill in the BLANC: Human-free quality estimation of document summaries (2020)
- Does GPT-2 know your phone number? (2020)
- Asking and answering questions to evaluate the factual consistency of summaries (2020)
- Continuity of topic, interaction, and query: Learning to quote in online conversations (2020)
- Applying the transformer to character-level transduction (2020)
- Recipes for safety in open-domain chatbots, 2020a (2020)
- AutoQA: From databases to QA semantic parsers with only synthetic training data (2020)
- ReClor: A reading comprehension dataset requiring logical reasoning (2020)
- Figure me out: A gold standard dataset for metaphor interpretation (2020)
- The gap of semantic parsing: A survey on automatic math word problem solvers (2020)
- Hurtful words: Quantifying biases in clinical contextual word embeddings (2020)
- WinoWhy: A deep diagnosis of essential commonsense knowledge for answering Winograd schema challenge (2020)
- Reasoning about goals, steps, and temporal ordering with WikiHow (2020)
- When do you need billions of words of pretraining data?, 2020e (2020)
- Detecting hallucinated content in conditional neural sequence generation (2020)
- Asking clarifying questions in open-domain information-seeking conversations (2019)
- MathQA: Towards interpretable math word problem solving with operation-based formalisms (2019)
- Toward automated quest generation in text-adventure games (2019)
- Big BiRD: A large, fine-grained, bigram relatedness dataset for examining semantic composition (2019)
- Real or fake? Learning to discriminate machine from human generated text (2019)
- Neural path planning: Fixed time, near-optimal path generation via oracle imitation (2019)
- Abductive commonsense reasoning (2019)
- Large dataset and language model fun-tuning for humor recognition (2019)
- Vector forms as a foreign language, 24 June 2019 (2019)
- Identifying and reducing gender bias in word-level language models (2019)
- COMET: Commonsense transformers for automatic knowledge graph construction (2019)
- Finding microaggressions in the wild: A case for locating elusive phenomena in social media posts (2019)
- Studying cultural differences in emoji usage across the East and the West (2019)
- Touchdown: Natural language navigation and spatial reasoning in visual street environments (2019)
- Execution-guided neural program synthesis (2019)
- Generating long sequences with sparse transformers (2019)
- On measuring gender bias in translation of gender-neutral pronouns (2019)
- On the measure of intelligence (2019)
- BoolQ: Exploring the surprising difficulty of natural yes/no questions (2019)
- The CommitmentBank: Investigating projection in naturally occurring discourse (2019)
- Queens are powerful too: Mitigating gender bias in dialogue generation (2019)
- DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs (2019)
- Misspelling oblivious word embeddings (2019)
- Parsimonious morpheme segmentation with an application to enriching word embeddings (2019)
- Making sense of sensory input (2019)
- Question answering as an automatic evaluation metric for news article summarization (2019)
- Teaching GPT-2 transformer a sense of humor: How to fine-tune large transformer models on a single GPU in PyTorch (2019)
- Knowledge-aware assessment of severity of suicide risk for early intervention (2019)
- Assessing BERT's syntactic abilities (2019)
- Lipstick on a pig: Debiasing methods cover up systematic gender biases in word embeddings but do not remove them (2019)
- Topical-chat: Towards knowledge-grounded open-domain conversations (2019)
- Stochastic optimization of sorting networks via continuous relaxations (2019)
- It's all in the name: Mitigating gender bias with name-based counterfactual data substitution (2019)
- Using pre-training can improve model robustness and uncertainty (2019)
- Beyond human-level accuracy: Computational challenges in deep learning (2019)
- National name report 2018 (2019)
- GQA: A new dataset for real-world visual reasoning and compositional question answering (2019)
- Do you know that Florence is packed with visitors? Evaluating state-of-the-art models of speaker commitment (2019)
- Rogue-Gym: A new challenge for generalization in reinforcement learning (2019)
- Negated and misprimed probes for pretrained language models: Birds can talk, but cannot fly (2019)
- Learning the difference that makes a difference with counterfactually-augmented data (2019)
- Cooperation and codenames: Understanding natural language processing via codenames (2019)
- A surprisingly robust trick for the Winograd schema challenge (2019)
- Natural questions: A benchmark for question answering research (2019)
- The emergence of number and syntax units in LSTM language models (2019)
- Deep learning for symbolic mathematics (2019)
- ALBERT: A lite BERT for self-supervised learning of language representations (2019)
- Revisiting the evaluation of theory of mind through question answering (2019)
- Reasoning over paragraph effects in situations (2019)
- RoBERTa: A robustly optimized BERT pretraining approach (2019)
- A survey of reinforcement learning informed by natural language (2019)
- Encode, tag, realize: High-precision text editing (2019)
- A BERT-based approach for automatic humor detection and scoring (2019)
- Suicide risk assessment with multi-level dual-context language and BERT (2019)
- On measuring social biases in sentence encoders (2019)
- Extending machine language models toward human-level language understanding (2019)
- Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference (2019)
- "Andere zeiten, andere lehren": Sprach-und kulturgeschichtliche betrachtungen zum sprichwort (2019)
- CLaC at CLPsych 2019: Fusion of neural features and predicted class probabilities for suicide risk assessment based on online posts (2019)
- More data can hurt for linear regression: Sample-wise double descent (2019)
- DisSent: Learning sentence representations from explicit discourse relations (2019)
- Generating natural anagrams: Towards language generation under hard combinatorial constraints (2019)
- Can you trust your model's uncertainty? Evaluating predictive uncertainty under dataset shift (2019)
- Learning algorithms via neural logic networks (2019)
- Deep and dense sarcasm detection (2019)
- Language models as knowledge bases? (2019)
- MELD: A multimodal multi-party dataset for emotion recognition in conversations (2019)
- Language models are unsupervised multitask learners (2019)
- Folk psychology as a theory (2019)
- CoQA: A conversational question answering challenge (2019)
- Sentence-BERT: Sentence embeddings using siamese BERT-networks (2019)
- A constructive prediction of the generalization error across scales (2019)
- How well do NLI models capture verb veridicality? (2019)
- Social IQa: Commonsense reasoning about social interactions (2019)
- Analysing mathematical reasoning abilities of neural models (2019)
- Language tasks and language games: On methodology in current natural language processing research (2019)
- Revisiting low-resource neural machine translation: A case study (2019)
- The woman worked as a babysitter: On biases in language generation (2019)
- Mining discourse markers for unsupervised sentence representation learning (2019)
- CLUTRR: A diagnostic benchmark for inductive reasoning from text (2019)
- Release strategies and the social impacts of language models (2019)
- Patching gender: Non-binary utopias in HCI (2019)
- Evaluating gender bias in machine translation (2019)
- Executing instructions in situated collaborative interactions (2019)
- The bitter lesson (2019)
- LSTM networks can perform dynamic counting (2019)
- Memory-augmented recurrent neural networks can learn generalized Dyck languages, 2019b (2019)
- oLMpics – on what language model pre-training captures, 2019a (2019)
- CommonsenseQA: A question answering challenge targeting commonsense knowledge (2019)
- The teaching size: Computable teachers and learners for universal languages (2019)
- Fine-grained temporal relation extraction (2019)
- Grandmaster level in StarCraft II using multi-agent reinforcement learning (2019)
- SuperGLUE: A stickier benchmark for general-purpose language understanding systems (2019)
- Learning to count objects with few exemplar annotations, 2019b (2019)
- SATNet: Bridging deep learning and logical reasoning using a differentiable satisfiability solver, 2019c (2019)
- TalkDown: A corpus for condescension detection in context (2019)
- Humor detection: A transformer gets the last laugh (2019)
- Some additional experiments extending the tech report "assessing BERT’s syntactic abilities" by Yoav Goldberg (2019)
- Huggingface's transformers: State-of-the-art natural language processing (2019)
- Learning to prove theorems via interacting with proof assistants (2019)
- CoSQL: A conversational text-to-SQL challenge towards cross-domain natural language interfaces to databases (2019)
- SParC: Cross-domain semantic parsing in context (2019)
- Learning the Dyck language with attention-based Seq2Seq models (2019)
- ActivityNet-QA: A dataset for understanding complex web videos via question answering, 2019d (2019)
- HellaSwag: Can a machine really finish your sentence? (2019)
- Defending against neural fake news, 2019b (2019)
- Multi-agent reinforcement learning: A selective overview of theories and algorithms, 2019a (2019)
- Irony detection via sentiment-based transfer learning (2019)
- "Going on a vacation" takes longer than "going for a walk": A study of temporal commonsense understanding (2019)
- Learning to ask unanswerable questions for machine reading comprehension (2019)
- A survey of machine learning for big code and naturalness (2018)
- code2seq: Generating sequences from structured representations of code (2018)
- Predicting human metaphor paraphrase judgments with deep neural networks (2018)
- The WMT'18 morpheval test suites for English-Czech, English-German, English-Finnish and Turkish-English (2018)
- Humor recognition using deep learning (2018)
- QuAC: Question answering in context (2018)
- Think you have solved question answering? Try ARC, the AI2 reasoning challenge (2018)
- General-purpose declarative inductive programming with domain-specific background knowledge for data wrangling automation (2018)
- Introduction to Logic (2018)
- Snips voice platform: An embedded spoken language understanding system for private-by-design voice interfaces (2018)
- TextWorld: A learning environment for text-based games (2018)
- BERT: Pre-training of deep bidirectional transformers for language understanding (2018)
- Compositional morpheme embeddings with affixes as functions and stems as arguments (2018)
- Semantic relatedness of Wikipedia concepts – benchmark data and a working solution (2018)
- Can neural networks understand logical entailment? (2018)
- Hierarchical neural story generation (2018)
- Whodunnit? Crime drama as a case for natural language understanding (2018)
- “The penny drops”: Investigating insight through the medium of cryptic crosswords (2018)
- Neural metaphor detection in context (2018)
- Universal neural machine translation for extremely low resource languages (2018)
- Colorless green recurrent networks dream hierarchically (2018)
- The argument reasoning comprehension task: Identification and reconstruction of implicit warrants (2018)
- Context-free transductions with neural stacks (2018)
- Women also snowboard: Overcoming bias in captioning models (2018)
- Algorithmic regulation and the rule of law (2018)
- Gamepad: A learning environment for theorem proving (2018)
- AI safety via debate (2018)
- The misgendering machines: Trans/HCI implications of automatic gender recognition (2018)
- Recurrent Neural Networks in Linguistic Theory: Revisiting Pinker and Prince (1988) and the Past Tense Debate (2018)
- The NarrativeQA reading comprehension challenge (2018)
- WikiHow: A large scale text summarization dataset (2018)
- SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing (2018)
- Scalable agent alignment via reward modeling: A research direction (2018)
- A meaning-based statistical English math word problem solver (2018)
- Content preserving text generation with attribute controls (2018)
- Automatic prediction of discourse connectives (2018)
- Targeted syntactic evaluation of language models (2018)
- The application of convolution neural network based cell segmentation during cryopreservation (2018)
- The natural language decathlon: Multitask learning as question answering (2018)
- Revisiting the poverty of the stimulus: Hierarchical generalization without a hierarchical bias in recurrent neural networks (2018)
- Interactive optimal teaching with unknown learners (2018)
- National name statistical analysis (2018)
- Stress test evaluation for natural language inference (2018)
- Don't give me the details, just the summary! Topic-aware convolutional neural networks for extreme summarization (2018)
- Evaluating theory of mind in question answering (2018)
- Know what you don't know: Unanswerable questions for SQuAD (2018)
- Gender bias in coreference resolution (2018)
- Learning a SAT solver from single-bit supervision (2018)
- Evaluating the ability of LSTMs to learn context-free grammars (2018)
- Adafactor: Adaptive learning rates with sublinear memory cost (2018)
- Expert, crowdsourced, and machine assessment of suicide risk via online postings (2018)
- Closing brackets with recurrent neural networks (2018)
- FEVER: A large-scale dataset for fact extraction and VERification (2018)
- Neural arithmetic logic units (2018)
- Dating documents using graph convolution networks (2018)
- GLUE: A multi-task benchmark and analysis platform for natural language understanding (2018)
- It's going to be okay: Measuring access to support in online communities (2018)
- Lexicosyntactic inference in neural models (2018)
- A broad-coverage challenge corpus for sentence understanding through inference (2018)
- Incorporating latent meanings of morphological compositions to enhance word embeddings (2018)
- Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task (2018)
- From recognition to cognition: Visual commonsense reasoning (2018)
- ReCoRD: Bridging the gap between human and machine commonsense reading comprehension, 2018a (2018)
- Learning to count objects in natural images for visual question answering, 2018b (2018)
- Gender bias in coreference resolution: Evaluation and debiasing methods (2018)
- Optnet: Differentiable optimization as a layer in neural networks (2017)
- Humor in language (2017)
- Neural-symbolic learning and reasoning: A survey and interpretation (2017)
- Deep API programmer: Learning to program with APIs (2017)
- Programming with a differentiable Forth interpreter (2017)
- Semantics derived automatically from language corpora contain human-like biases (2017)
- CycleGAN, a master of steganography (2017)
- The trouble with bias (2017)
- Language modeling with gated convolutional networks (2017)
- RobustFill: Neural program learning under noisy I/O (2017)
- Quasar: Datasets for question answering by search and reading (2017)
- Learning to learn programs from examples: Going beyond program structure (2017)
- Using millions of emoji occurrences to learn any-domain representations for detecting sentiment, emotion and sarcasm (2017)
- Color naming across languages reflects color use (2017)
- Program synthesis (2017)
- On calibration of modern neural networks (2017)
- A baseline for detecting misclassified and out-of-distribution examples in neural networks (2017)
- Deep learning scaling is predictable, empirically (2017)
- Smelling themselves: Dogs investigate their own odours longer when modified in an “olfactory mirror” test (2017)
- Can self-awareness be taught? Monkeys pass the mirror test – again (2017)
- Automatic sarcasm detection: A survey (2017)
- TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension (2017)
- A large self-annotated corpus for sarcasm (2017)
- Self-Aware Computing Systems (2017)
- RACE: Large-scale ReAding comprehension dataset from examinations (2017)
- Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks (2017)
- Building machines that learn and think like people (2017)
- Zero-shot relation extraction via reading comprehension (2017)
- DailyDialog: A manually labelled multi-turn dialogue dataset (2017)
- Program induction by rationale generation: Learning to solve and explain algebraic word problems (2017)
- Temporal information extraction for question answering using syntactic dependencies in an LSTM-based architecture (2017)
- SemEval-2017 task 7: Detection and interpretation of English puns (2017)
- Word sense disambiguation: A unified evaluation framework and empirical comparison (2017)
- Semi-supervised multitask learning for sequence labeling (2017)
- Neural joke generation (2017)
- Automatic detection of satire in Twitter: A psycholinguistic-based approach (2017)
- A simple neural network module for relational reasoning (2017)
- Cryopreservation aims to engineer novel ways to freeze, store, and thaw organs (2017)
- Get to the point: Summarization with pointer-generator networks (2017)
- Prerequisite skills for reading comprehension: Multi-perspective analysis of MCTest datasets and systems (2017)
- Attention is all you need (2017)
- Computational argumentation quality assessment in natural language (2017)
- Cognitive and emotional demands of black humour processing: The role of intelligence, aggressiveness and mood (2017)
- Who's to say what's funny? A computer using language models and deep learning, that's who! (2017)
- Explainable artificial intelligence via Bayesian teaching (2017)
- Learning continuous semantic representations of symbolic expressions (2016)
- Concrete problems in AI safety (2016)
- Deepcoder: Learning to write programs (2016)
- ITEM2VEC: Neural item embedding for collaborative filtering (2016)
- Big data's disparate impact (2016)
- Man is to computer programmer as woman is to homemaker? Debiasing word embeddings (2016)
- Ravens attribute visual access to unseen competitors (2016)
- Meta-interpretive learning of data transformation programs (2016)
- emoji2vec: Learning emoji representations from their description (2016)
- TerpreT: A probabilistic programming language for program induction (2016)
- Pragmatic language interpretation as probabilistic inference (2016)
- Hybrid computing using a neural network with dynamic external memory (2016)
- Tracking the world state with recurrent entity networks (2016)
- CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning (2016)
- Assessing the ability of LSTMs to learn syntax-sensitive dependencies (2016)
- Pointer sentinel mixture models (2016)
- Introducing the LCC metaphor datasets (2016)
- A corpus and evaluation framework for deeper understanding of commonsense stories (2016)
- Program synthesis from polymorphic refinement types (2016)
- What are the most popular names chinese parents give their babies? a perspective from big data (2016)
- SQuAD: 100,000+ questions for machine comprehension of text (2016)
- Robsut wrod reocginiton via semi-character recurrent neural network (2016)
- How grammatical is character-level neural machine translation? Assessing MT quality with contrastive translation pairs (2016)
- Transforming spreadsheet data types using examples (2016)
- Inferring interpersonal relations in narrative summaries (2016)
- Learning language games through interaction (2016)
- Google's neural machine translation system: Bridging the gap between human and machine translation (2016)
- Tweet sarcasm detection using deep neural network (2016)
- The secrets of colour interpolation, 6 Jan. 2016 (2016)
- VQA: Visual question answering (2015)
- A large annotated corpus for learning natural language inference (2015)
- Emoji usage in TV conversation (2015)
- Synthesizing data structure transformations from input-output examples (2015)
- Inductive programming meets the real world (2015)
- The MovieLens datasets: History and context (2015)
- Teaching machines to read and comprehend (2015)
- Introduction to Paremiology: A Comprehensive Guide to Proverb Studies (2015)
- Emojineering part 1: Machine learning for emoji trends (2015)
- Harnessing context incongruity for sarcasm detection (2015)
- Inferring algorithmic patterns with stack-augmented recurrent nets (2015)
- The unreasonable effectiveness of recurrent neural networks (2015)
- Character-aware neural language models (2015)
- TR9856: A multi-word term relatedness benchmark (2015)
- SemEval-2015 task 5: QA TempEval - evaluating temporal information understanding with question answering (2015)
- Annotating character relationships in literary texts (2015)
- Image-based recommendations on styles and substitutes (2015)
- Automatic disambiguation of English puns (2015)
- Obtaining well calibrated probabilities using Bayesian binning (2015)
- Posterior calibration and exploratory analysis for natural language processing models (2015)
- Standard Occupational Classification of the People's Republic of China (2015)
- Type-and-example-directed program synthesis (2015)
- SemEval 2015, task 7: Diachronic text evaluation (2015)
- Neural programmer-interpreters (2015)
- Number-space mapping in the newborn chick resembles humans’ mental number line (2015)
- A neural attention model for abstractive sentence summarization (2015)
- Neural machine translation of rare words with subword units (2015)
- Predicting a correct program in programming by example (2015)
- Learning to recommend quotes for writing (2015)
- Learning to interpret natural language commands through human-robot dialog (2015)
- Survey on collaborative filtering, content-based filtering and hybrid recommendation system (2015)
- Towards AI-complete question answering: A set of prerequisite toy tasks (2015)
- Humor recognition and humor anchor extraction (2015)
- Optimizing sentence modeling and selection for document summarization (2015)
- Machine teaching: An inverse problem to machine learning and an approach toward optimal education (2015)
- Non-linear mapping for improved identification of 1300+ languages (2014)
- Eliciting good teaching from humans for machine learners (2014)
- Sequence-based prediction of protein–protein interaction sites with l1-logreg classifier (2014)
- The Cattell-Horn-Carroll theory of cognitive abilities (2014)
- The Elements of Eloquence: Secrets of the Perfect Turn of Phrase (2014)
- Dr.Fill: Crosswords and an implemented solver for singly weighted CSPs (2014)
- Neural Turing machines (2014)
- Learning to solve arithmetic word problems with verb categorization (2014)
- A burstiness-aware approach for document dating (2014)
- Diagram understanding in geometry questions (2014)
- SPRINGS: Prediction of protein-protein interaction sites using artificial neural networks (2014)
- Sequence to sequence learning with neural networks (2014)
- Metaphor detection with cross-lingual model transfer (2014)
- Learning to execute (2014)
- Teaching classification boundaries to humans (2013)
- Rosetta stone linguistic problems (2013)
- Semantic sort: A supervised approach to personalized semantic relatedness (2013)
- Global inference for bridging anaphora resolution (2013)
- Generating expressions that refer to visible objects (2013)
- Playing Atari with deep reinforcement learning (2013)
- "The things that we have to do": Ethics and instrumentality in humanitarian communication (2013)
- Recursive deep models for semantic compositionality over a sentiment treebank (2013)
- Labeling documents with timestamps: Learning from their time expressions (2012)
- Did it happen? The pragmatic complexity of veridicality assessment (2012)
- Spreadsheet data manipulation using examples (2012)
- Analogy and relational reasoning (2012)
- Roget's Thesaurus as a lexical resource for natural language processing (2012)
- Simple and phrasal implicatives (2012)
- Collective classification for fine-grained information status (2012)
- Resolving complex cases of definite pronouns: The Winograd schema challenge (2012)
- Learning data transformation rules through examples: Preliminary results (2012)
- D^3 data-driven documents (2011)
- Identifying sarcasm in Twitter: A closer look (2011)
- Automating string processing in spreadsheets using input-output examples (2011)
- Wrangler: Interactive visual specification of data transformation scripts (2011)
- How do humans teach: On curriculum learning and teaching dimension (2011)
- A word at a time: Computing word relatedness using temporal semantic analysis (2011)
- Choice of plausible alternatives: An evaluation of commonsense causal reasoning (2011)
- Who uses web search for what: And how (2011)
- Machine translation of Klingon (2010)
- Choice of plausible alternatives (COPA) (2010)
- The weirdest people in the world? (2010)
- Inductive programming: A survey of program synthesis techniques (2010)
- Natural reference to objects in a visual domain (2010)
- Applying the naïve Bayes classifier with kernel density estimation to the prediction of protein–protein interaction sites (2010)
- Multi-prototype vector-space models of word meaning (2010)
- Temporal reasoning in natural language processing: A survey (2010)
- Automatic metaphor interpretation as a paraphrasing task (2010)
- Metaphor corpus annotated for source-target domain mappings (2010)
- A Method for Linguistic Metaphor Identification: From MIP to MIPVU (2010)
- Text relatedness based on a word thesaurus (2010)
- The TUNA-REG challenge 2009: Overview and evaluation results (2009)
- Sprichwörter und zweisprachige lexikographie: Deutsch-schwedische und deutsch-finnische wörtebücher im vergleich (2009)
- The aha! moment: The cognitive neuroscience of insight (2009)
- Omiotis: A thesaurus-based measure of text relatedness (2009)
- Revisiting the strange stories: Revealing mentalizing impairments in autism (2009)
- WikiWalk: Random walks on Wikipedia for semantic relatedness (2009)
- Does the chimpanzee have a theory of mind? 30 years later (2008)
- Finding contradictions in text (2008)
- An introduction to inductive programming (2008)
- Metaphors We Live By (2008)
- An effective, low-cost measure of semantic relatedness obtained from Wikipedia links (2008)
- The North American computational linguistics olympiad (NACLO) (2008)
- The New York Times annotated corpus LDC2008T19 (2008)
- The use of spatial relations in referring expression generation (2008)
- Wechsler Adult Intelligence Scale–Fourth Edition (WAIS–IV) (2008)
- Artificial intelligence as a positive and negative factor in global risk (2008)
- Russian Proverbs and Sayings and Their English Equivalents (2007)
- The Google similarity distance (2007)
- Computing semantic relatedness using Wikipedia-based explicit semantic analysis (2007)
- When in Rome, do as the Romans do: Proverbs as a part of EFL teaching (2007)
- Lexical semantic relatedness with random graph walks (2007)
- Phraseologie des schwedischen (2007)
- Comparisons of sequence labeling algorithms and extensions (2007)
- Knowledge derived from Wikipedia for computing semantic relatedness (2007)
- A clustering approach for nearly unsupervised recognition of nonliteral language (2006)
- Wikirelate! Computing semantic relatedness using Wikipedia (2006)
- Making computers laugh: Investigations in automatic humor recognition (2005)
- The Montreal Cognitive Assessment, MoCA: A brief screening tool for mild cognitive impairment (2005)
- Effects of directionality in deductive reasoning, II. Premise integration and conclusion evaluation (2005)
- Causation and causal inference in epidemiology (2005)
- Authorship verification as a one-class classification problem (2004)
- Solving logic puzzles: From robust processing to precise semantics (2004)
- The specification language TimeML (2004)
- Extended gloss overlaps as a measure of semantic relatedness (2003)
- Simplicity: A unifying principle in cognitive science? (2003)
- English gigaword (2003)
- Intentional action and side effects in ordinary language (2003)
- An approach for measuring semantic similarity between words using multiple information sources (2003)
- Holographic Reduced Representations: Distributed Representation for Cognitive Structures (2003)
- Revisions that improve cohesion in multi-document summaries: A preliminary study (2002)
- Artificial Intelligence: A Modern Approach (2002)
- Pragmatics, modularity and mind-reading (2002)
- Neural networks – a model of boolean functions (2002)
- Proverb comprehension as a function of reading proficiency in preadolescents (2001)
- Effects of directionality in deductive reasoning, I. The comprehension of single relational premises (2000)
- Causality: Models, Reasoning, and Inference (2000)
- Causation, Prediction, and Search (2000)
- Inference making ability and its relation to comprehension failure (1999)
- Semantic similarity in a taxonomy: An information-based measure and its application to problems of ambiguity in natural language (1999)
- WordNet: An Electronic Lexical Database (1998)
- A Proverb in Mind: The Cognitive Science of Proverbial Wit and Wisdom (1997)
- Semantic similarity based on corpus statistics and lexical taxonomy (1997)
- The proverb as a mitigating and politeness strategy in Akan discourse (1996)
- Celex2 ldc96l14 (1995)
- Using information content to evaluate semantic similarity in a taxonomy (1995)
- Self recognition in a jumping spider: Portia labiata females discriminate between their own draglines and those of conspecifics (1994)
- An advanced test of theory of mind: Understanding of story characters thoughts and feelings by able autistic, mentally handicapped, and normal children and adults (1994)
- The Penn Treebank: Annotating predicate argument structure (1994)
- Distributed representations and nested compositional structure (1994)
- Controlling other people: The impact of power on stereotyping (1993)
- The roles of similarity in transfer: Separating retrievability from inferential soundness (1993)
- Context based spelling correction (1991)
- Lexical cohesion computed by thesaural relations as an indicator of the structure of text (1991)
- Indexing by latent semantic analysis (1990)
- Analog retrieval by constraint satisfaction (1990)
- They Never Said It: A Book of Fake Quotes, Misquotes, and Misleading Attributions (1989)
- Connectionism and cognitive architecture: A critical analysis (1988)
- Comprehending complex concepts (1988)
- History of Kannada Literature: Readership Lectures (1988)
- Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference (1988)
- On the proper treatment of connectionism (1988)
- Parallel Distributed Processing. Volume 1: Foundations (1986)
- The synthesis of LISP programs from examples: A survey (1984)
- On the projection problem for presuppositions (1983)
- Application of theorem proving to problem solving (1981)
- A general psychoevolutionary theory of emotion (1980)
- The inference of regular LISP programs from examples (1978)
- Killing, letting die, and the trolley problem (1976)
- The Language of Thought (1975)
- Joking riddles: A developmental index of children's humor (1975)
- Inferring LISP programs from examples (1975)
- Progress report on program-understanding systems (AIM-240) (1974)
- More is different (1972)
- Structural equation methods in the social sciences (1972)
- Understanding natural language (1972)
- English Proverbs and Sayings (1971)
- On a family of Turing machines and the related programming language (1964)
- The algebraic theory of context-free languages (1959)
- The child's learning of english morphology (1958)
- In defense of a dogma (1956)
- Main trends in recent philosophy: Two dogmas of empiricism (1951)
- Verification of forecasts expressed in terms of probability (1950)
- Computing machinery and intelligence (1950)
- Prolegomena to Any Future Metaphysics (1783)
- Critique of Pure Reason (1781)
- A Treatise of Human Nature (1739)
- Wikiquote, russian proverbs