User profiles for Todor Mihaylov
Todor MihaylovResearch Scientist at Meta Superintelligence Labs Verified email at meta.com Cited by 58255 |
Filtering, distillation, and hard negatives for vision-language pre-training
Vision-language models trained with contrastive learning on large-scale noisy data are
becoming increasingly popular for zero-shot recognition problems. In this paper we improve the …
becoming increasingly popular for zero-shot recognition problems. In this paper we improve the …
Llama 2: Open foundation and fine-tuned chat models
In this work, we develop and release Llama 2, a collection of pretrained and fine-tuned large
language models (LLMs) ranging in scale from 7 billion to 70 billion parameters. Our fine-…
language models (LLMs) ranging in scale from 7 billion to 70 billion parameters. Our fine-…
Opt: Open pre-trained transformer language models
Large language models, which are often trained for hundreds of thousands of compute days,
have shown remarkable capabilities for zero- and few-shot learning. Given their …
have shown remarkable capabilities for zero- and few-shot learning. Given their …
Few-shot learning with multilingual generative language models
Large-scale generative language models such as GPT-3 are competitive few-shot learners.
While these models are known to be able to jointly represent many different languages, their …
While these models are known to be able to jointly represent many different languages, their …
Can a suit of armor conduct electricity? a new dataset for open book question answering
We present a new kind of question answering dataset, OpenBookQA, modeled after open
book exams for assessing human understanding of a subject. The open book that comes with …
book exams for assessing human understanding of a subject. The open book that comes with …
[PDF][PDF] Finding opinion manipulation trolls in news community forums
T Mihaylov, G Georgiev, P Nakov - Proceedings of the nineteenth …, 2015 - aclanthology.org
The emergence of user forums in electronic news media has given rise to the proliferation of
opinion manipulation trolls. Finding such trolls automatically is a hard task, as there is no …
opinion manipulation trolls. Finding such trolls automatically is a hard task, as there is no …
Knowledgeable reader: Enhancing cloze-style reading comprehension with external commonsense knowledge
T Mihaylov, A Frank - Proceedings of the 56th Annual Meeting of …, 2018 - aclanthology.org
We introduce a neural reading comprehension model that integrates external commonsense
knowledge, encoded as a key-value memory, in a cloze-style setting. Instead of relying only …
knowledge, encoded as a key-value memory, in a cloze-style setting. Instead of relying only …
Opt-iml: Scaling language model instruction meta learning through the lens of generalization
Recent work has shown that fine-tuning large pre-trained language models on a collection
of tasks described via instructions, aka instruction-tuning, improves their zero and few-shot …
of tasks described via instructions, aka instruction-tuning, improves their zero and few-shot …
Efficient large scale language modeling with mixtures of experts
Mixture of Experts layers (MoEs) enable efficient scaling of language models through
conditional computation. This paper presents a detailed empirical study of how autoregressive …
conditional computation. This paper presents a detailed empirical study of how autoregressive …
[PDF][PDF] Hunting for troll comments in news community forums
T Mihaylov, P Nakov - Proceedings of the 54th Annual Meeting of …, 2016 - aclanthology.org
There are different definitions of what a troll is. Certainly, a troll can be somebody who teases
people to make them angry, or somebody who offends people, or somebody who wants to …
people to make them angry, or somebody who offends people, or somebody who wants to …