Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–15 of 15 results for author: Golub, M

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.03893  [pdf, ps, other

    cs.LG

    Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse

    Authors: Taekyung Heo, Rasoul Shafipour, Ritchie Zhao, Maximilian Golub, Mohammad Mahdi Kamani, Ritika Borkar, Makesh Tarun Chandran, Pantea Zardoshti, Bita Darvish Rouhani

    Abstract: Production deployments often swap between different-sized models in a family for cost-quality cascading, mid-conversation switching, and routing, and each swap forces the receiver to repay the prefill from scratch. We propose cross-model KV cache transfer, where the receiver reuses the source's KV cache, skipping prefill. We find that cross-model KV has substantial linear structure across matched-… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

  2. arXiv:2606.15007  [pdf, ps, other

    cs.CL cs.AI cs.LG

    Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

    Authors: NVIDIA, :, Aaron Blakeman, Aaron Thomas, Aastha Jhunjhunwala, Abhibha Gupta, Abhinav Khattar, Adam Rajfer, Adi Renduchintala, Adil Asif, Aditya Vavre, Adriana Flores Miranda, Ahmad Bilal, Aileen Zaman, Ajay Hotchandani, Akanksha Shukla, Akhiad Bercovich, Aleksander Ficek, Alex Gronskiy, Alex Kondratenko, Alex Steiner, Alex Ye, Alexander Bukharin, Alexandre Milesi, Ali Taghibakhshi , et al. (549 additional authors not shown)

    Abstract: We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model. We pre-trained Nemotron 3 Ultra on 20 trillion text tokens, then extended the context length to 1M tokens, and post-trained using Supervised Fine Tuning (SFT), Reinforcement Learning (RL), and Multi-teacher On-Policy Distillation (MOPD). Nemotron 3 Ultra is o… ▽ More

    Submitted 12 June, 2026; originally announced June 2026.

  3. arXiv:2601.18089  [pdf, ps, other

    cs.LG cs.AI

    LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts

    Authors: Venmugil Elango, Nidhi Bhatia, Roger Waleffe, Rasoul Shafipour, Tomer Asida, Abhinav Khattar, Nave Assaf, Maximilian Golub, Joey Guman, Tiyasa Mitra, Ritchie Zhao, Ritika Borkar, Ran Zilberstein, Mostofa Patwary, Mohammad Shoeybi, Bita Rouhani

    Abstract: Mixture of Experts (MoEs) have become a central component of many state-of-the-art open-source and proprietary large language models. Despite their widespread adoption, it remains unclear how close existing MoE architectures are to optimal with respect to inference cost, as measured by accuracy per floating-point operation and per parameter. In this work, we revisit MoE design from a hardware-soft… ▽ More

    Submitted 25 January, 2026; originally announced January 2026.

  4. arXiv:2507.07120  [pdf, ps, other

    cs.DC cs.AI

    Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding

    Authors: Nidhi Bhatia, Ankit More, Ritika Borkar, Tiyasa Mitra, Ramon Matas, Ritchie Zhao, Maximilian Golub, Dheevatsa Mudigere, Brian Pharris, Bita Darvish Rouhani

    Abstract: As LLMs scale to multi-million-token KV histories, real-time autoregressive decoding under tight Token-to-Token Latency (TTL) constraints faces growing pressure. Two core bottlenecks dominate: accessing Feed-Forward Network (FFN) weights and reading long KV caches. While Tensor Parallelism (TP) helps mitigate the cost of FFN weight reads, it does not scale well for attention. When TP width exceeds… ▽ More

    Submitted 7 July, 2025; originally announced July 2025.

  5. arXiv:2506.19094  [pdf, ps, other

    q-bio.NC cs.CE

    Accurate identification of communication between multiple interacting neural populations

    Authors: Belle Liu, Jacob Sacks, Matthew D. Golub

    Abstract: Neural recording technologies now enable simultaneous recording of population activity across many brain regions, motivating the development of data-driven models of inter-regional communication. However, existing models can struggle to disentangle the influences that drive recorded population activity, leading to inaccurate portraits of communication. Here, we introduce Multi-Region Latent Factor… ▽ More

    Submitted 5 June, 2026; v1 submitted 23 June, 2025; originally announced June 2025.

    Journal ref: Forty-second International Conference on Machine Learning (2025)

  6. arXiv:2506.05508  [pdf, ps, other

    cs.DC cs.AI

    Beyond the Buzz: A Pragmatic Take on Inference Disaggregation

    Authors: Tiyasa Mitra, Ritika Borkar, Nidhi Bhatia, Ramon Matas, Shivam Raj, Dheevatsa Mudigere, Ritchie Zhao, Maximilian Golub, Arpan Dutta, Sailaja Madduri, Dharmesh Jani, Brian Pharris, Bita Darvish Rouhani

    Abstract: As inference scales to multi-node deployments, disaggregation - splitting inference into distinct phases - offers a promising path to improving the throughput-interactivity Pareto frontier. Despite growing enthusiasm and a surge of open-source efforts, practical deployment of disaggregated serving remains limited due to the complexity of the optimization search space and system-level coordination.… ▽ More

    Submitted 5 June, 2025; originally announced June 2025.

  7. arXiv:2412.02529  [pdf, other

    q-bio.NC cs.LG stat.ML

    Active learning of neural population dynamics using two-photon holographic optogenetics

    Authors: Andrew Wagenmaker, Lu Mi, Marton Rozsa, Matthew S. Bull, Karel Svoboda, Kayvon Daie, Matthew D. Golub, Kevin Jamieson

    Abstract: Recent advances in techniques for monitoring and perturbing neural populations have greatly enhanced our ability to study circuits in the brain. In particular, two-photon holographic optogenetics now enables precise photostimulation of experimenter-specified groups of individual neurons, while simultaneous two-photon calcium imaging enables the measurement of ongoing and induced activity across th… ▽ More

    Submitted 8 May, 2025; v1 submitted 3 December, 2024; originally announced December 2024.

    Comments: NeurIPS 2024

  8. arXiv:2402.11723  [pdf, other

    cs.HC cs.CL

    Shaping Human-AI Collaboration: Varied Scaffolding Levels in Co-writing with Language Models

    Authors: Paramveer S. Dhillon, Somayeh Molaei, Jiaqi Li, Maximilian Golub, Shaochun Zheng, Lionel P. Robert

    Abstract: Advances in language modeling have paved the way for novel human-AI co-writing experiences. This paper explores how varying levels of scaffolding from large language models (LLMs) shape the co-writing process. Employing a within-subjects field experiment with a Latin square design, we asked participants (N=131) to respond to argumentative writing prompts under three randomly sequenced conditions:… ▽ More

    Submitted 18 February, 2024; originally announced February 2024.

    Comments: Appearing at CHI 2024 (Honolulu, HI)

  9. arXiv:2310.10537  [pdf, other

    cs.LG cs.AI

    Microscaling Data Formats for Deep Learning

    Authors: Bita Darvish Rouhani, Ritchie Zhao, Ankit More, Mathew Hall, Alireza Khodamoradi, Summer Deng, Dhruv Choudhary, Marius Cornea, Eric Dellinger, Kristof Denolf, Stosic Dusan, Venmugil Elango, Maximilian Golub, Alexander Heinecke, Phil James-Roxby, Dharmesh Jani, Gaurav Kolhe, Martin Langhammer, Ada Li, Levi Melnick, Maral Mesmakhosroshahi, Andres Rodriguez, Michael Schulte, Rasoul Shafipour, Lei Shao , et al. (8 additional authors not shown)

    Abstract: Narrow bit-width data formats are key to reducing the computational and storage costs of modern deep learning applications. This paper evaluates Microscaling (MX) data formats that combine a per-block scaling factor with narrow floating-point and integer types for individual elements. MX formats balance the competing needs of hardware efficiency, model accuracy, and user friction. Empirical result… ▽ More

    Submitted 19 October, 2023; v1 submitted 16 October, 2023; originally announced October 2023.

  10. arXiv:2302.08007  [pdf, other

    cs.LG cs.AI cs.AR

    With Shared Microexponents, A Little Shifting Goes a Long Way

    Authors: Bita Rouhani, Ritchie Zhao, Venmugil Elango, Rasoul Shafipour, Mathew Hall, Maral Mesmakhosroshahi, Ankit More, Levi Melnick, Maximilian Golub, Girish Varatkar, Lei Shao, Gaurav Kolhe, Dimitry Melts, Jasmine Klar, Renee L'Heureux, Matt Perry, Doug Burger, Eric Chung, Zhaoxia Deng, Sam Naghshineh, Jongsoo Park, Maxim Naumov

    Abstract: This paper introduces Block Data Representations (BDR), a framework for exploring and evaluating a wide spectrum of narrow-precision formats for deep learning. It enables comparison of popular quantization standards, and through BDR, new formats based on shared microexponents (MX) are identified, which outperform other state-of-the-art quantization approaches, including narrow-precision floating-p… ▽ More

    Submitted 12 April, 2023; v1 submitted 15 February, 2023; originally announced February 2023.

  11. Enriched physics-informed neural networks for in-plane crack problems: Theory and MATLAB codes

    Authors: Yan Gu, Chuanzeng Zhang, Peijun Zhang, Mikhail V. Golub

    Abstract: In this paper, a method based on the physics-informed neural networks (PINNs) is presented to model in-plane crack problems in the linear elastic fracture mechanics. Instead of forming a mesh, the PINNs is meshless and can be trained on batches of randomly sampled collocation points. In order to capture the theoretical singular behavior of the near-tip stress and strain fields, the standard PINNs… ▽ More

    Submitted 12 June, 2022; originally announced June 2022.

    Comments: 31 pages, 17 figures, 5 tables

    MSC Class: 35M32

    Journal ref: International Journal of Solids and Structures, Volume 276, 1 August 2023, 112321

  12. arXiv:2009.10976  [pdf, other

    cs.NE cs.AR cs.LG

    Procrustes: a Dataflow and Accelerator for Sparse Deep Neural Network Training

    Authors: Dingqing Yang, Amin Ghasemazar, Xiaowei Ren, Maximilian Golub, Guy Lemieux, Mieszko Lis

    Abstract: The success of DNN pruning has led to the development of energy-efficient inference accelerators that support pruned models with sparse weight and activation tensors. Because the memory layouts and dataflows in these architectures are optimized for the access patterns during $\mathit{inference}$, however, they do not efficiently support the emerging sparse $\mathit{training}$ techniques. In this… ▽ More

    Submitted 23 September, 2020; originally announced September 2020.

    Comments: Appears in the Proceedings of the 53$^\mathit{rd}$ IEEE/ACM International Symposium on Microarchitecture (MICRO 2020)

  13. arXiv:1907.08549  [pdf, other

    q-bio.NC cs.NE

    Universality and individuality in neural dynamics across large populations of recurrent networks

    Authors: Niru Maheswaranathan, Alex H. Williams, Matthew D. Golub, Surya Ganguli, David Sussillo

    Abstract: Task-based modeling with recurrent neural networks (RNNs) has emerged as a popular way to infer the computational function of different brain regions. These models are quantitatively assessed by comparing the low-dimensional neural representations of the model with the brain, for example using canonical correlation analysis (CCA). However, the nature of the detailed neurobiological inferences one… ▽ More

    Submitted 4 December, 2019; v1 submitted 19 July, 2019; originally announced July 2019.

    Comments: Presented at NeurIPS 2019

  14. arXiv:1906.10720  [pdf, other

    cs.LG stat.ML

    Reverse engineering recurrent networks for sentiment classification reveals line attractor dynamics

    Authors: Niru Maheswaranathan, Alex Williams, Matthew D. Golub, Surya Ganguli, David Sussillo

    Abstract: Recurrent neural networks (RNNs) are a widely used tool for modeling sequential data, yet they are often treated as inscrutable black boxes. Given a trained recurrent network, we would like to reverse engineer it--to obtain a quantitative, interpretable description of how it solves a particular task. Even for simple tasks, a detailed understanding of how recurrent networks work, or a prescription… ▽ More

    Submitted 4 December, 2019; v1 submitted 25 June, 2019; originally announced June 2019.

    Comments: Presented at NeurIPS 2019

  15. arXiv:1806.06949  [pdf, other

    cs.LG stat.ML

    Full deep neural network training on a pruned weight budget

    Authors: Maximilian Golub, Guy Lemieux, Mieszko Lis

    Abstract: We introduce a DNN training technique that learns only a fraction of the full parameter set without incurring an accuracy penalty. To do this, our algorithm constrains the total number of weights updated during backpropagation to those with the highest total gradients. The remaining weights are not tracked, and their initial value is regenerated at every access to avoid storing them in memory. Thi… ▽ More

    Submitted 23 November, 2019; v1 submitted 11 June, 2018; originally announced June 2018.

    Journal ref: Proceedings of the 2nd SysML Conference, Palo Alto, CA, USA, 2019