Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–35 of 35 results for author: Michel, P

Searching in archive cs. Search in all archives.
.
  1. arXiv:2507.06261  [pdf, ps, other

    cs.CL cs.AI

    Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

    Authors: Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, Luke Marris, Sam Petulla, Colin Gaffney, Asaf Aharoni, Nathan Lintz, Tiago Cardal Pais, Henrik Jacobsson, Idan Szpektor, Nan-Jiang Jiang, Krishna Haridasan, Ahmed Omran, Nikunj Saunshi, Dara Bahri, Gaurav Mishra, Eric Chu , et al. (3410 additional authors not shown)

    Abstract: In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our most capable model yet, achieving SoTA performance on frontier coding and reasoning benchmarks. In addition to its incredible coding and reasoning skills, Gemini 2.5 Pro is a thinking model that excels at multimodal unde… ▽ More

    Submitted 19 December, 2025; v1 submitted 7 July, 2025; originally announced July 2025.

    Comments: 72 pages, 17 figures

  2. arXiv:2504.20563  [pdf, other

    cs.LO

    Turing machines deciders, part I

    Authors: The bbchallenge Collaboration, Justin Blanchard, Konrad Deka, Nathan Fenner, Tony Guilfoyle, Iijil, Maja Kądziołka, Pavel Kropitz, Shawn Ligocki, Pascal Michel, Mateusz Naściszewski, Tristan Stérin

    Abstract: The Busy Beaver Challenge (or bbchallenge) aims at collaboratively solving the following conjecture: "$S(5) = 47{,}176{,}870$" [Radó, 1962], [Marxen and Buntrock, 1990], [Aaronson, 2020]. This conjecture says that if a 5-state Turing machine runs for more than 47,176,870 steps without halting, then it will never halt -- starting from the all-0 tape. Proving this conjecture amounts to deciding whet… ▽ More

    Submitted 29 April, 2025; originally announced April 2025.

    Comments: 41 pages

    ACM Class: F.1; F.1.1; F.1.3; F.4.1

  3. A large collection of bioinformatics question-query pairs over federated knowledge graphs: methodology and applications

    Authors: Jerven Bolleman, Vincent Emonet, Adrian Altenhoff, Amos Bairoch, Marie-Claude Blatter, Alan Bridge, Severine Duvaud, Elisabeth Gasteiger, Dmitry Kuznetsov, Sebastien Moretti, Pierre-Andre Michel, Anne Morgat, Marco Pagni, Nicole Redaschi, Monique Zahn-Zabal, Tarcisio Mendes de Farias, Ana Claudia Sima

    Abstract: Background. In the last decades, several life science resources have structured data using the same framework and made these accessible using the same query language to facilitate interoperability. Knowledge graphs have seen increased adoption in bioinformatics due to their advantages for representing data in a generic graph format. For example, yummydata.org catalogs more than 60 knowledge graphs… ▽ More

    Submitted 8 October, 2024; originally announced October 2024.

    Journal ref: GigaScience, Volume 14, 2025

  4. arXiv:2408.00118  [pdf, other

    cs.CL cs.AI

    Gemma 2: Improving Open Language Models at a Practical Size

    Authors: Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan, Sammy Jerome, Anton Tsitsulin, Nino Vieillard, Piotr Stanczyk, Sertan Girgin, Nikola Momchev, Matt Hoffman , et al. (173 additional authors not shown)

    Abstract: In this work, we introduce Gemma 2, a new addition to the Gemma family of lightweight, state-of-the-art open models, ranging in scale from 2 billion to 27 billion parameters. In this new version, we apply several known technical modifications to the Transformer architecture, such as interleaving local-global attentions (Beltagy et al., 2020a) and group-query attention (Ainslie et al., 2023). We al… ▽ More

    Submitted 2 October, 2024; v1 submitted 31 July, 2024; originally announced August 2024.

  5. arXiv:2406.11409  [pdf, other

    cs.CL cs.AI

    CodeGemma: Open Code Models Based on Gemma

    Authors: CodeGemma Team, Heri Zhao, Jeffrey Hui, Joshua Howland, Nam Nguyen, Siqi Zuo, Andrea Hu, Christopher A. Choquette-Choo, Jingyue Shen, Joe Kelley, Kshitij Bansal, Luke Vilnis, Mateo Wirth, Paul Michel, Peter Choy, Pratik Joshi, Ravin Kumar, Sarmad Hashmi, Shubham Agrawal, Zhitao Gong, Jane Fine, Tris Warkentin, Ale Jakse Hartman, Bin Ni, Kathy Korevec , et al. (2 additional authors not shown)

    Abstract: This paper introduces CodeGemma, a collection of specialized open code models built on top of Gemma, capable of a variety of code and natural language generation tasks. We release three model variants. CodeGemma 7B pretrained (PT) and instruction-tuned (IT) variants have remarkably resilient natural language understanding, excel in mathematical reasoning, and match code capabilities of other open… ▽ More

    Submitted 18 June, 2024; v1 submitted 17 June, 2024; originally announced June 2024.

    Comments: v1: 11 pages, 4 figures, 5 tables. v2: Update metadata

  6. arXiv:2404.19409  [pdf, other

    cs.CL

    Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning

    Authors: Mathieu Rita, Florian Strub, Rahma Chaabouni, Paul Michel, Emmanuel Dupoux, Olivier Pietquin

    Abstract: While Reinforcement Learning (RL) has been proven essential for tuning large language models (LLMs), it can lead to reward over-optimization (ROO). Existing approaches address ROO by adding KL regularization, requiring computationally expensive hyperparameter tuning. Additionally, KL regularization focuses solely on regularizing the language policy, neglecting a potential source of regularization:… ▽ More

    Submitted 30 April, 2024; originally announced April 2024.

  7. arXiv:2403.11958  [pdf, other

    cs.CL cs.MA

    Language Evolution with Deep Learning

    Authors: Mathieu Rita, Paul Michel, Rahma Chaabouni, Olivier Pietquin, Emmanuel Dupoux, Florian Strub

    Abstract: Computational modeling plays an essential role in the study of language emergence. It aims to simulate the conditions and learning processes that could trigger the emergence of a structured language within a simulated controlled environment. Several methods have been used to investigate the origin of our language, including agent-based systems, Bayesian agents, genetic algorithms, and rule-based s… ▽ More

    Submitted 18 March, 2024; originally announced March 2024.

    Comments: to appear in the Oxford Handbook of Approaches to Language Evolution

  8. arXiv:2403.08295  [pdf, other

    cs.CL cs.AI

    Gemma: Open Models Based on Gemini Research and Technology

    Authors: Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambrose Slone, Amélie Héliou, Andrea Tacchetti, Anna Bulanova, Antonia Paterson, Beth Tsai, Bobak Shahriari , et al. (83 additional authors not shown)

    Abstract: This work introduces Gemma, a family of lightweight, state-of-the art open models built from the research and technology used to create Gemini models. Gemma models demonstrate strong performance across academic benchmarks for language understanding, reasoning, and safety. We release two sizes of models (2 billion and 7 billion parameters), and provide both pretrained and fine-tuned checkpoints. Ge… ▽ More

    Submitted 16 April, 2024; v1 submitted 13 March, 2024; originally announced March 2024.

  9. arXiv:2403.05530  [pdf, other

    cs.CL cs.AI

    Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

    Authors: Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, Soroosh Mariooryad, Yifan Ding, Xinyang Geng, Fred Alcober, Roy Frostig, Mark Omernick, Lexi Walker, Cosmin Paduraru, Christina Sorokin, Andrea Tacchetti, Colin Gaffney, Samira Daruki, Olcan Sercinoglu, Zach Gleicher, Juliette Love , et al. (1112 additional authors not shown)

    Abstract: In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over fine-grained information from millions of tokens of context, including multiple long documents and hours of video and audio. The family includes two new models: (1) an updated Gemini 1.5 Pro, which exceeds the February… ▽ More

    Submitted 16 December, 2024; v1 submitted 8 March, 2024; originally announced March 2024.

  10. arXiv:2312.11805  [pdf, other

    cs.CL cs.AI cs.CV

    Gemini: A Family of Highly Capable Multimodal Models

    Authors: Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, Melvin Johnson, Ioannis Antonoglou, Julian Schrittwieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy Lillicrap, Angeliki Lazaridou, Orhan Firat, James Molloy, Michael Isard, Paul R. Barham, Tom Hennigan, Benjamin Lee , et al. (1326 additional authors not shown)

    Abstract: This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consists of Ultra, Pro, and Nano sizes, suitable for applications ranging from complex reasoning tasks to on-device memory-constrained use-cases. Evaluation on a broad range of benchmarks shows that our most-capable Gemini Ultr… ▽ More

    Submitted 9 May, 2025; v1 submitted 18 December, 2023; originally announced December 2023.

  11. arXiv:2308.12202  [pdf, other

    cs.LG cs.CL

    Curriculum Learning with Adam: The Devil Is in the Wrong Details

    Authors: Lucas Weber, Jaap Jumelet, Paul Michel, Elia Bruni, Dieuwke Hupkes

    Abstract: Curriculum learning (CL) posits that machine learning models -- similar to humans -- may learn more efficiently from data that match their current learning progress. However, CL methods are still poorly understood and, in particular for natural language processing (NLP), have achieved only limited success. In this paper, we explore why. Starting from an attempt to replicate and extend a number of… ▽ More

    Submitted 23 August, 2023; originally announced August 2023.

  12. arXiv:2209.15342  [pdf, other

    cs.MA cs.CL cs.IT

    Emergent Communication: Generalization and Overfitting in Lewis Games

    Authors: Mathieu Rita, Corentin Tallec, Paul Michel, Jean-Bastien Grill, Olivier Pietquin, Emmanuel Dupoux, Florian Strub

    Abstract: Lewis signaling games are a class of simple communication games for simulating the emergence of language. In these games, two agents must agree on a communication protocol in order to solve a cooperative task. Previous work has shown that agents trained to play this game with reinforcement learning tend to develop languages that display undesirable properties from a linguistic point of view (lack… ▽ More

    Submitted 15 October, 2022; v1 submitted 30 September, 2022; originally announced September 2022.

    Comments: 36th Conference on Neural Information Processing Systems (NeurIPS 2022)

  13. arXiv:2209.04751  [pdf, other

    cs.RO

    Enabling Under-Ice Geochemical Observations with a Size, Weight, and Power-Constrained Robot

    Authors: Jess Horowitz, Victoria Preston, Anna P. M. Michel

    Abstract: Estimates of greenhouse gas emissions from Arctic estuarine environments are dominated by in situ summer-time ice-free dissolved gas measurements due to the logistical ease of performing field observations in these conditions. Recent evidence in coastal Arctic environments, however, has demonstrated that dissolved methane (CH4) and carbon dioxide (CO2) are strongly seasonally variable, and at leas… ▽ More

    Submitted 10 September, 2022; originally announced September 2022.

    Comments: 6 pages, 7 figures, OCEANS Conference 2022

  14. arXiv:2206.01364  [pdf, other

    cs.RO

    Robotic Planning under Uncertainty in Spatiotemporal Environments in Expeditionary Science

    Authors: Victoria Preston, Genevieve Flaspohler, Anna P. M. Michel, John W. Fisher III, Nicholas Roy

    Abstract: In the expeditionary sciences, spatiotemporally varying environments -- hydrothermal plumes, algal blooms, lava flows, or animal migrations -- are ubiquitous. Mobile robots are uniquely well-suited to study these dynamic, mesoscale natural environments. We formalize expeditionary science as a sequential decision-making problem, modeled using the language of partially-observable Markov decision pro… ▽ More

    Submitted 2 June, 2022; originally announced June 2022.

    Comments: 5 pages, 1 figure, as submitted to The Multi-disciplinary Conference on Reinforcement Learning and Decision Making

  15. arXiv:2205.14082  [pdf, other

    cs.LG cs.AI

    AANG: Automating Auxiliary Learning

    Authors: Lucio M. Dery, Paul Michel, Mikhail Khodak, Graham Neubig, Ameet Talwalkar

    Abstract: Auxiliary objectives, supplementary learning signals that are introduced to help aid learning on data-starved or highly complex end-tasks, are commonplace in machine learning. Whilst much work has been done to formulate useful auxiliary objectives, their construction is still an art which proceeds by slow and tedious hand-design. Intuition for how and when these objectives improve end-task perform… ▽ More

    Submitted 27 February, 2023; v1 submitted 27 May, 2022; originally announced May 2022.

    Comments: Accepted to ICLR 2023 22 pages, 7 tables and 5 figures

  16. arXiv:2204.06340  [pdf, other

    cs.LG

    Distributionally Robust Models with Parametric Likelihood Ratios

    Authors: Paul Michel, Tatsunori Hashimoto, Graham Neubig

    Abstract: As machine learning models are deployed ever more broadly, it becomes increasingly important that they are not only able to perform well on their training distribution, but also yield accurate predictions when confronted with distribution shift. The Distributionally Robust Optimization (DRO) framework proposes to address this issue by training models to minimize their expected risk under a collect… ▽ More

    Submitted 13 April, 2022; originally announced April 2022.

    Comments: ICLR 2022

  17. arXiv:2110.05838  [pdf, other

    cs.LG cs.AI cs.CL

    Balancing Average and Worst-case Accuracy in Multitask Learning

    Authors: Paul Michel, Sebastian Ruder, Dani Yogatama

    Abstract: When training and evaluating machine learning models on a large number of tasks, it is important to not only look at average task accuracy -- which may be biased by easy or redundant tasks -- but also worst-case accuracy (i.e. the performance on the task with the lowest accuracy). In this work, we show how to use techniques from the distributionally robust optimization (DRO) literature to improve… ▽ More

    Submitted 12 October, 2021; originally announced October 2021.

    Comments: Under review

  18. arXiv:2109.07437  [pdf, other

    cs.LG cs.CL

    Should We Be Pre-training? An Argument for End-task Aware Training as an Alternative

    Authors: Lucio M. Dery, Paul Michel, Ameet Talwalkar, Graham Neubig

    Abstract: In most settings of practical concern, machine learning practitioners know in advance what end-task they wish to boost with auxiliary tasks. However, widely used methods for leveraging auxiliary data like pre-training and its continued-pretraining variant are end-task agnostic: they rarely, if ever, exploit knowledge of the target task. We study replacing end-task agnostic continued training of pr… ▽ More

    Submitted 6 February, 2022; v1 submitted 15 September, 2021; originally announced September 2021.

    Comments: 18 pages, 4 figures

  19. arXiv:2109.01558  [pdf, other

    cs.CL

    Learning Neural Models for Natural Language Processing in the Face of Distributional Shift

    Authors: Paul Michel

    Abstract: The dominating NLP paradigm of training a strong neural predictor to perform one task on a specific dataset has led to state-of-the-art performance in a variety of applications (eg. sentiment classification, span-prediction based question answering or machine translation). However, it builds upon the assumption that the data distribution is stationary, ie. that the data is sampled from a fixed dis… ▽ More

    Submitted 3 September, 2021; originally announced September 2021.

    Comments: PhD thesis

    Report number: CMU-LTI-21-013

  20. arXiv:2106.07171  [pdf, other

    cs.LG cs.AI

    Examining and Combating Spurious Features under Distribution Shift

    Authors: Chunting Zhou, Xuezhe Ma, Paul Michel, Graham Neubig

    Abstract: A central goal of machine learning is to learn robust representations that capture the causal relationship between inputs features and output labels. However, minimizing empirical risk over finite or biased datasets often results in models latching on to spurious correlations between the training input/output pairs that are not fundamental to the problem at hand. In this paper, we define and analy… ▽ More

    Submitted 14 June, 2021; originally announced June 2021.

    Comments: Accepted by ICML2021

  21. arXiv:2103.10282  [pdf, other

    cs.LG cs.CL

    Modeling the Second Player in Distributionally Robust Optimization

    Authors: Paul Michel, Tatsunori Hashimoto, Graham Neubig

    Abstract: Distributionally robust optimization (DRO) provides a framework for training machine learning models that are able to perform well on a collection of related data distributions (the "uncertainty set"). This is done by solving a min-max game: the model is trained to minimize its maximum expected loss among all distributions in the uncertainty set. While careful design of the uncertainty set is crit… ▽ More

    Submitted 31 March, 2021; v1 submitted 18 March, 2021; originally announced March 2021.

    Comments: Accepted at ICLR 2021

  22. arXiv:2005.00783  [pdf, other

    cs.LG cs.CR stat.ML

    Differentially Private Generation of Small Images

    Authors: Justus T. C. Schwabedal, Pascal Michel, Mario S. Riontino

    Abstract: We explore the training of generative adversarial networks with differential privacy to anonymize image data sets. On MNIST, we numerically measure the privacy-utility trade-off using parameters from $ε$-$δ$ differential privacy and the inception score. Our experiments uncover a saturated training regime where an increasing privacy budget adds little to the quality of generated images. We also exp… ▽ More

    Submitted 6 May, 2020; v1 submitted 2 May, 2020; originally announced May 2020.

    Comments: 11 pages, 3 figures. Revised criticism of Beaulieu-Jones et al (2017): their results are likely correct

  23. arXiv:2004.06660  [pdf, other

    cs.LG cs.CL cs.CR stat.ML

    Weight Poisoning Attacks on Pre-trained Models

    Authors: Keita Kurita, Paul Michel, Graham Neubig

    Abstract: Recently, NLP has seen a surge in the usage of large pre-trained models. Users download weights of models pre-trained on large datasets, then fine-tune the weights on a task of their choice. This raises the question of whether downloading untrusted pre-trained weights can pose a security threat. In this paper, we show that it is possible to construct ``weight poisoning'' attacks where pre-trained… ▽ More

    Submitted 14 April, 2020; originally announced April 2020.

    Comments: Published as a long paper at ACL 2020

  24. arXiv:1911.10088  [pdf, other

    cs.LG cs.CL stat.ML

    Optimizing Data Usage via Differentiable Rewards

    Authors: Xinyi Wang, Hieu Pham, Paul Michel, Antonios Anastasopoulos, Jaime Carbonell, Graham Neubig

    Abstract: To acquire a new skill, humans learn better and faster if a tutor, based on their current knowledge level, informs them of how much attention they should pay to particular content or practice problems. Similarly, a machine learning model could potentially be trained better with a scorer that "adapts" to its current learning state and estimates the importance of each training data instance. Trainin… ▽ More

    Submitted 16 June, 2021; v1 submitted 22 November, 2019; originally announced November 2019.

    Comments: Accepted at ICML 2020

  25. Information-Guided Robotic Maximum Seek-and-Sample in Partially Observable Continuous Environments

    Authors: Genevieve Flaspohler, Victoria Preston, Anna P. M. Michel, Yogesh Girdhar, Nicholas Roy

    Abstract: We present PLUMES, a planner to localizing and collecting samples at the global maximum of an a priori unknown and partially observable continuous environment. The "maximum-seek-and-sample" (MSS) problem is pervasive in the environmental and earth sciences. Experts want to collect scientifically valuable samples at an environmental maximum (e.g., an oil-spill source), but do not have prior knowled… ▽ More

    Submitted 26 September, 2019; originally announced September 2019.

    Comments: 8 pages, 8 figures, To appear in the proceedings of IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2019 Macau

    Journal ref: IEEE Robotics and Automation Letters (RA-L) 2019

  26. arXiv:1906.11943  [pdf, other

    cs.CL

    Findings of the First Shared Task on Machine Translation Robustness

    Authors: Xian Li, Paul Michel, Antonios Anastasopoulos, Yonatan Belinkov, Nadir Durrani, Orhan Firat, Philipp Koehn, Graham Neubig, Juan Pino, Hassan Sajjad

    Abstract: We share the findings of the first shared task on improving robustness of Machine Translation (MT). The task provides a testbed representing challenges facing MT models deployed in the real world, and facilitates new approaches to improve models; robustness to noisy input and domain mismatch. We focus on two language pairs (English-French and English-Japanese), and the submitted systems are evalua… ▽ More

    Submitted 3 July, 2019; v1 submitted 27 June, 2019; originally announced June 2019.

  27. arXiv:1905.10650  [pdf, other

    cs.CL

    Are Sixteen Heads Really Better than One?

    Authors: Paul Michel, Omer Levy, Graham Neubig

    Abstract: Attention is a powerful and ubiquitous mechanism for allowing neural models to focus on particular salient pieces of information by taking their weighted average when making predictions. In particular, multi-headed attention is a driving force behind many recent state-of-the-art NLP models such as Transformer-based MT models and BERT. These models apply multiple attention mechanisms in parallel, w… ▽ More

    Submitted 4 November, 2019; v1 submitted 25 May, 2019; originally announced May 2019.

    Comments: NeurIPS 2019

  28. arXiv:1903.07926  [pdf, other

    cs.CL

    compare-mt: A Tool for Holistic Comparison of Language Generation Systems

    Authors: Graham Neubig, Zi-Yi Dou, Junjie Hu, Paul Michel, Danish Pruthi, Xinyi Wang, John Wieting

    Abstract: In this paper, we describe compare-mt, a tool for holistic analysis and comparison of the results of systems for language generation tasks such as machine translation. The main goal of the tool is to give the user a high-level and coherent view of the salient differences between systems that can then be used to guide further analysis or system improvement. It implements a number of tools to do so,… ▽ More

    Submitted 19 September, 2019; v1 submitted 19 March, 2019; originally announced March 2019.

    Comments: Updated and longer version of NAACL 2019 Demo Paper

  29. arXiv:1903.06620  [pdf, other

    cs.CL

    On Evaluation of Adversarial Perturbations for Sequence-to-Sequence Models

    Authors: Paul Michel, Xian Li, Graham Neubig, Juan Miguel Pino

    Abstract: Adversarial examples --- perturbations to the input of a model that elicit large changes in the output --- have been shown to be an effective way of assessing the robustness of sequence-to-sequence (seq2seq) models. However, these perturbations only indicate weaknesses in the model if they do not change the input so significantly that it legitimately results in changes in the expected output. This… ▽ More

    Submitted 18 March, 2019; v1 submitted 15 March, 2019; originally announced March 2019.

    Comments: NAACL-HLT 2019 long paper

  30. arXiv:1809.00388  [pdf, other

    cs.CL

    MTNT: A Testbed for Machine Translation of Noisy Text

    Authors: Paul Michel, Graham Neubig

    Abstract: Noisy or non-standard input text can cause disastrous mistranslations in most modern Machine Translation (MT) systems, and there has been growing research interest in creating noise-robust MT systems. However, as of yet there are no publicly available parallel corpora of with naturally occurring noisy inputs and translations, and thus previous work has resorted to evaluating on synthetically creat… ▽ More

    Submitted 2 September, 2018; originally announced September 2018.

    Comments: EMNLP 2018 Long Paper

  31. arXiv:1805.01817  [pdf, other

    cs.CL

    Extreme Adaptation for Personalized Neural Machine Translation

    Authors: Paul Michel, Graham Neubig

    Abstract: Every person speaks or writes their own flavor of their native language, influenced by a number of factors: the content they tend to talk about, their gender, their social status, or their geographical origin. When attempting to perform Machine Translation (MT), these variations have a significant effect on how the system should perform translation, but this is not captured well by standard one-… ▽ More

    Submitted 4 May, 2018; originally announced May 2018.

    Comments: Accepted as a short paper at ACL 2018

  32. arXiv:1705.10900  [pdf, other

    cs.CL

    Does the Geometry of Word Embeddings Help Document Classification? A Case Study on Persistent Homology Based Representations

    Authors: Paul Michel, Abhilasha Ravichander, Shruti Rijhwani

    Abstract: We investigate the pertinence of methods from algebraic topology for text data analysis. These methods enable the development of mathematically-principled isometric-invariant mappings from a set of vectors to a document embedding, which is stable with respect to the geometry of the document in the selected metric space. In this work, we evaluate the utility of these topology-based document represe… ▽ More

    Submitted 30 May, 2017; originally announced May 2017.

    Comments: 5 pages, 3 figures. Rep4NLP workshop at ACL 2017

  33. arXiv:1701.03980  [pdf, other

    stat.ML cs.CL cs.MS

    DyNet: The Dynamic Neural Network Toolkit

    Authors: Graham Neubig, Chris Dyer, Yoav Goldberg, Austin Matthews, Waleed Ammar, Antonios Anastasopoulos, Miguel Ballesteros, David Chiang, Daniel Clothiaux, Trevor Cohn, Kevin Duh, Manaal Faruqui, Cynthia Gan, Dan Garrette, Yangfeng Ji, Lingpeng Kong, Adhiguna Kuncoro, Gaurav Kumar, Chaitanya Malaviya, Paul Michel, Yusuke Oda, Matthew Richardson, Naomi Saphra, Swabha Swayamdipta, Pengcheng Yin

    Abstract: We describe DyNet, a toolkit for implementing neural network models based on dynamic declaration of network structure. In the static declaration strategy that is used in toolkits like Theano, CNTK, and TensorFlow, the user first defines a computation graph (a symbolic representation of the computation), and then examples are fed into an engine that executes this computation and computes its deriva… ▽ More

    Submitted 14 January, 2017; originally announced January 2017.

    Comments: 33 pages

  34. arXiv:1608.00508  [pdf, other

    cs.CL

    Blind phoneme segmentation with temporal prediction errors

    Authors: Paul Michel, Okko Räsänen, Roland Thiollière, Emmanuel Dupoux

    Abstract: Phonemic segmentation of speech is a critical step of speech recognition systems. We propose a novel unsupervised algorithm based on sequence prediction models such as Markov chains and recurrent neural network. Our approach consists in analyzing the error profile of a model trained to predict speech features frame-by-frame. Specifically, we try to learn the dynamics of speech in the MFCC space an… ▽ More

    Submitted 27 May, 2017; v1 submitted 1 August, 2016; originally announced August 2016.

    Comments: 7 pages 3 figures. Presented at ACL SRW 2017

  35. arXiv:1311.1029  [pdf, ps, other

    math.LO cs.CC cs.LO

    Problems in number theory from busy beaver competition

    Authors: Pascal Michel

    Abstract: By introducing the busy beaver competition of Turing machines, in 1962, Rado defined noncomputable functions on positive integers. The study of these functions and variants leads to many mathematical challenges. This article takes up the following one: How can a small Turing machine manage to produce very big numbers? It provides the following answer: mostly by simulating Collatz-like functions,… ▽ More

    Submitted 11 December, 2015; v1 submitted 5 November, 2013; originally announced November 2013.

    Comments: 35 pages

    Journal ref: Logical Methods in Computer Science, Volume 11, Issue 4 (December 14, 2015) lmcs:1611