User profiles for Todor Mihaylov

Todor Mihaylov

Research Scientist at Meta Superintelligence Labs
Verified email at meta.com
Cited by 58255

Filtering, distillation, and hard negatives for vision-language pre-training

…, A Dubey, A Kadian, T Mihaylov… - 2023 IEEE/CVF …, 2023 - ieeexplore.ieee.org
Vision-language models trained with contrastive learning on large-scale noisy data are
becoming increasingly popular for zero-shot recognition problems. In this paper we improve the …

Llama 2: Open foundation and fine-tuned chat models

…, Y Lu, Y Mao, X Martinet, T Mihaylov… - arXiv preprint arXiv …, 2023 - arxiv.org
In this work, we develop and release Llama 2, a collection of pretrained and fine-tuned large
language models (LLMs) ranging in scale from 7 billion to 70 billion parameters. Our fine-…

Opt: Open pre-trained transformer language models

…, C Dewan, M Diab, X Li, XV Lin, T Mihaylov… - arXiv preprint arXiv …, 2022 - arxiv.org
Large language models, which are often trained for hundreds of thousands of compute days,
have shown remarkable capabilities for zero- and few-shot learning. Given their …

Few-shot learning with multilingual generative language models

XV Lin, T Mihaylov, M Artetxe, T Wang… - Proceedings of the …, 2022 - aclanthology.org
Large-scale generative language models such as GPT-3 are competitive few-shot learners.
While these models are known to be able to jointly represent many different languages, their …

Can a suit of armor conduct electricity? a new dataset for open book question answering

T Mihaylov, P Clark, T Khot… - Proceedings of the 2018 …, 2018 - aclanthology.org
We present a new kind of question answering dataset, OpenBookQA, modeled after open
book exams for assessing human understanding of a subject. The open book that comes with …

[PDF][PDF] Finding opinion manipulation trolls in news community forums

T Mihaylov, G Georgiev, P Nakov - Proceedings of the nineteenth …, 2015 - aclanthology.org
The emergence of user forums in electronic news media has given rise to the proliferation of
opinion manipulation trolls. Finding such trolls automatically is a hard task, as there is no …

Knowledgeable reader: Enhancing cloze-style reading comprehension with external commonsense knowledge

T Mihaylov, A Frank - Proceedings of the 56th Annual Meeting of …, 2018 - aclanthology.org
We introduce a neural reading comprehension model that integrates external commonsense
knowledge, encoded as a key-value memory, in a cloze-style setting. Instead of relying only …

Opt-iml: Scaling language model instruction meta learning through the lens of generalization

S Iyer, XV Lin, R Pasunuru, T Mihaylov, D Simig… - arXiv preprint arXiv …, 2022 - arxiv.org
Recent work has shown that fine-tuning large pre-trained language models on a collection
of tasks described via instructions, aka instruction-tuning, improves their zero and few-shot …

Efficient large scale language modeling with mixtures of experts

M Artetxe, S Bhosale, N Goyal, T Mihaylov… - arXiv preprint arXiv …, 2021 - arxiv.org
Mixture of Experts layers (MoEs) enable efficient scaling of language models through
conditional computation. This paper presents a detailed empirical study of how autoregressive …

[PDF][PDF] Hunting for troll comments in news community forums

T Mihaylov, P Nakov - Proceedings of the 54th Annual Meeting of …, 2016 - aclanthology.org
There are different definitions of what a troll is. Certainly, a troll can be somebody who teases
people to make them angry, or somebody who offends people, or somebody who wants to …