Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–16 of 16 results for author: Lam, T K

Searching in archive cs. Search in all archives.
.
  1. arXiv:2601.11329  [pdf, ps, other

    cs.CL

    F-Actor: Controllable Conversational Behaviour in Full-Duplex Models

    Authors: Maike Züfle, Ondrej Klejch, Nicholas Sanders, Jan Niehues, Alexandra Birch, Tsz Kin Lam

    Abstract: Spoken conversational systems require more than accurate speech generation to have human-like conversations: to feel natural and engaging, they must produce conversational behaviour that adapts dynamically to the context. Current spoken conversational systems, however, rarely allow such customization, limiting their naturalness and usability. In this work, we present the first open, instruction-fo… ▽ More

    Submitted 15 April, 2026; v1 submitted 16 January, 2026; originally announced January 2026.

  2. arXiv:2508.00537  [pdf, ps, other

    cs.CL

    The Prosody of Emojis

    Authors: Giulio Zhou, Tsz Kin Lam, Alexandra Birch, Barry Haddow

    Abstract: Prosodic features such as pitch, timing, and intonation are central to spoken communication, conveying emotion, intent, and discourse structure. In text-based settings, where these cues are absent, emojis act as visual surrogates that add affective and pragmatic nuance. This study examines how emojis influence prosodic realisation in speech and how listeners interpret prosodic cues to recover emoj… ▽ More

    Submitted 4 June, 2026; v1 submitted 1 August, 2025; originally announced August 2025.

    Comments: ACL 26

  3. arXiv:2503.10620  [pdf, ps, other

    cs.CL

    From TOWER to SPIRE: Adding the Speech Modality to a Translation-Specialist LLM

    Authors: Kshitij Ambilduke, Ben Peters, Sonal Sannigrahi, Anil Keshwani, Tsz Kin Lam, Bruno Martins, André F. T. Martins, Marcely Zanon Boito

    Abstract: We introduce Spire, a speech-augmented language model (LM) capable of both translating and transcribing speech input from English into 10 other languages as well as translating text input in both language directions. Spire integrates the speech modality into an existing multilingual LM via speech discretization and continued pre-training using only 42.5K hours of speech. In particular, we adopt th… ▽ More

    Submitted 22 October, 2025; v1 submitted 13 March, 2025; originally announced March 2025.

    Comments: EMNLP 2025 (Findings) camera ready

  4. arXiv:2501.02370  [pdf, other

    cs.CL cs.SD eess.AS

    Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison

    Authors: Tsz Kin Lam, Marco Gaido, Sara Papi, Luisa Bentivogli, Barry Haddow

    Abstract: Following the remarkable success of Large Language Models (LLMs) in NLP tasks, there is increasing interest in extending their capabilities to speech -- the most common form of communication. The most widespread approach to integrating speech into LLMs is dense feature prepending (DFP), which prepends the projected speech representations to the textual representations, allowing end-to-end training… ▽ More

    Submitted 7 February, 2025; v1 submitted 4 January, 2025; originally announced January 2025.

    Comments: Accepted at NAACL 2025

  5. arXiv:2411.05088  [pdf

    cs.CL

    Findings of the IWSLT 2024 Evaluation Campaign

    Authors: Ibrahim Said Ahmad, Antonios Anastasopoulos, Ondřej Bojar, Claudia Borg, Marine Carpuat, Roldano Cattoni, Mauro Cettolo, William Chen, Qianqian Dong, Marcello Federico, Barry Haddow, Dávid Javorský, Mateusz Krubiński, Tsz Kin Lam, Xutai Ma, Prashant Mathur, Evgeny Matusov, Chandresh Maurya, John McCrae, Kenton Murray, Satoshi Nakamura, Matteo Negri, Jan Niehues, Xing Niu, Atul Kr. Ojha , et al. (20 additional authors not shown)

    Abstract: This paper reports on the shared tasks organized by the 21st IWSLT Conference. The shared tasks address 7 scientific challenges in spoken language translation: simultaneous and offline translation, automatic subtitling and dubbing, speech-to-speech translation, dialect and low-resource speech translation, and Indic languages. The shared tasks attracted 18 teams whose submissions are documented in… ▽ More

    Submitted 7 November, 2024; originally announced November 2024.

    Comments: IWSLT 2024; 59 pages

  6. arXiv:2408.15366  [pdf, other

    cs.CL

    Pitfalls and Outlooks in Using COMET

    Authors: Vilém Zouhar, Pinzhen Chen, Tsz Kin Lam, Nikita Moghe, Barry Haddow

    Abstract: The COMET metric has blazed a trail in the machine translation community, given its strong correlation with human judgements of translation quality. Its success stems from being a modified pre-trained multilingual model finetuned for quality assessment. However, it being a machine learning model also gives rise to a new set of pitfalls that may not be widely known. We investigate these unexpected… ▽ More

    Submitted 30 September, 2024; v1 submitted 27 August, 2024; originally announced August 2024.

  7. arXiv:2402.19333  [pdf, other

    cs.CL cs.SD eess.AS

    Compact Speech Translation Models via Discrete Speech Units Pretraining

    Authors: Tsz Kin Lam, Alexandra Birch, Barry Haddow

    Abstract: We propose a pretraining method to use Self-Supervised Speech (SSS) model to creating more compact Speech-to-text Translation. In contrast to using the SSS model for initialization, our method is more suitable to memory constrained scenario such as on-device deployment. Our method is based on Discrete Speech Units (DSU) extracted from the SSS model. In the first step, our method pretrains two smal… ▽ More

    Submitted 26 June, 2024; v1 submitted 29 February, 2024; originally announced February 2024.

    Comments: 11 pages, accepted at IWSLT 2024

  8. arXiv:2402.00632  [pdf, other

    cs.CL

    Prosody in Cascade and Direct Speech-to-Text Translation: a case study on Korean Wh-Phrases

    Authors: Giulio Zhou, Tsz Kin Lam, Alexandra Birch, Barry Haddow

    Abstract: Speech-to-Text Translation (S2TT) has typically been addressed with cascade systems, where speech recognition systems generate a transcription that is subsequently passed to a translation model. While there has been a growing interest in developing direct speech translation systems to avoid propagating errors and losing non-verbal content, prior work in direct S2TT has struggled to conclusively es… ▽ More

    Submitted 1 February, 2024; originally announced February 2024.

    Comments: Accepted at Findings of EACL 2024

  9. arXiv:2302.03839  [pdf, other

    eess.IV cs.CV cs.LG

    Futuristic Variations and Analysis in Fundus Images Corresponding to Biological Traits

    Authors: Muhammad Hassan, Hao Zhang, Ahmed Fateh Ameen, Home Wu Zeng, Shuye Ma, Wen Liang, Dingqi Shang, Jiaming Ding, Ziheng Zhan, Tsz Kwan Lam, Ming Xu, Qiming Huang, Dongmei Wu, Can Yang Zhang, Zhou You, Awiwu Ain, Pei Wu Qin

    Abstract: Fundus image captures rear of an eye, and which has been studied for the diseases identification, classification, segmentation, generation, and biological traits association using handcrafted, conventional, and deep learning methods. In biological traits estimation, most of the studies have been carried out for the age prediction and gender classification with convincing results. However, the curr… ▽ More

    Submitted 7 February, 2023; originally announced February 2023.

    Comments: 10 pages, 4 figures, 3 tables

  10. Make More of Your Data: Minimal Effort Data Augmentation for Automatic Speech Recognition and Translation

    Authors: Tsz Kin Lam, Shigehiko Schamoni, Stefan Riezler

    Abstract: Data augmentation is a technique to generate new training data based on existing data. We evaluate the simple and cost-effective method of concatenating the original data examples to build new training instances. Continued training with such augmented data is able to improve off-the-shelf Transformer and Conformer models that were optimized on the original data only. We demonstrate considerable im… ▽ More

    Submitted 14 April, 2023; v1 submitted 27 October, 2022; originally announced October 2022.

    Comments: Accepted at ICASSP 2023

  11. arXiv:2210.13281  [pdf, other

    cs.CL

    Analyzing the Use of Influence Functions for Instance-Specific Data Filtering in Neural Machine Translation

    Authors: Tsz Kin Lam, Eva Hasler, Felix Hieber

    Abstract: Customer feedback can be an important signal for improving commercial machine translation systems. One solution for fixing specific translation errors is to remove the related erroneous training instances followed by re-training of the machine translation system, which we refer to as instance-specific data filtering. Influence functions (IF) have been shown to be effective in finding such relevant… ▽ More

    Submitted 24 October, 2022; originally announced October 2022.

    Comments: Accepted at WMT 2022

  12. Sample, Translate, Recombine: Leveraging Audio Alignments for Data Augmentation in End-to-end Speech Translation

    Authors: Tsz Kin Lam, Shigehiko Schamoni, Stefan Riezler

    Abstract: End-to-end speech translation relies on data that pair source-language speech inputs with corresponding translations into a target language. Such data are notoriously scarce, making synthetic data augmentation by back-translation or knowledge distillation a necessary ingredient of end-to-end training. In this paper, we present a novel approach to data augmentation that leverages audio alignments,… ▽ More

    Submitted 16 March, 2022; originally announced March 2022.

    Comments: Accepted at ACL 2022

  13. On-the-Fly Aligned Data Augmentation for Sequence-to-Sequence ASR

    Authors: Tsz Kin Lam, Mayumi Ohta, Shigehiko Schamoni, Stefan Riezler

    Abstract: We propose an on-the-fly data augmentation method for automatic speech recognition (ASR) that uses alignment information to generate effective training samples. Our method, called Aligned Data Augmentation (ADA) for ASR, replaces transcribed tokens and the speech representations in an aligned manner to generate previously unseen training pairs. The speech representations are sampled from an audio… ▽ More

    Submitted 9 June, 2021; v1 submitted 3 April, 2021; originally announced April 2021.

    Comments: Accepted at INTERSPEECH 2021

  14. Cascaded Models With Cyclic Feedback For Direct Speech Translation

    Authors: Tsz Kin Lam, Shigehiko Schamoni, Stefan Riezler

    Abstract: Direct speech translation describes a scenario where only speech inputs and corresponding translations are available. Such data are notoriously limited. We present a technique that allows cascades of automatic speech recognition (ASR) and machine translation (MT) to exploit in-domain direct speech translation data in addition to out-of-domain MT and ASR data. After pre-training MT and ASR, we use… ▽ More

    Submitted 11 February, 2021; v1 submitted 21 October, 2020; originally announced October 2020.

    Comments: Accepted at ICASSP 2021

  15. arXiv:1907.02326  [pdf, other

    cs.CL

    Interactive-Predictive Neural Machine Translation through Reinforcement and Imitation

    Authors: Tsz Kin Lam, Shigehiko Schamoni, Stefan Riezler

    Abstract: We propose an interactive-predictive neural machine translation framework for easier model personalization using reinforcement and imitation learning. During the interactive translation process, the user is asked for feedback on uncertain locations identified by the system. Responses are weak feedback in the form of "keep" and "delete" edits, and expert demonstrations in the form of "substitute" e… ▽ More

    Submitted 5 July, 2019; v1 submitted 4 July, 2019; originally announced July 2019.

    Comments: Machine Translation Summit 2019 (MTSUMMIT XVII), Dublin, Ireland

  16. arXiv:1805.01553  [pdf, other

    cs.CL stat.ML

    A Reinforcement Learning Approach to Interactive-Predictive Neural Machine Translation

    Authors: Tsz Kin Lam, Julia Kreutzer, Stefan Riezler

    Abstract: We present an approach to interactive-predictive neural machine translation that attempts to reduce human effort from three directions: Firstly, instead of requiring humans to select, correct, or delete segments, we employ the idea of learning from human reinforcements in form of judgments on the quality of partial translations. Secondly, human effort is further reduced by using the entropy of wor… ▽ More

    Submitted 5 June, 2018; v1 submitted 3 May, 2018; originally announced May 2018.

    Comments: Published at EAMT 2018; Updated algorithm