Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–2 of 2 results for author: Rao, K U M

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.00029  [pdf, ps, other

    cs.AI cs.AR cs.LG cs.PL

    Nova: An End-to-End MLIR Compiler for Deep Learning

    Authors: Adwaid Suresh, Aparna A, Harshini V M, Jona Delcy C A, Killi Uma Maheswara Rao, Ram Charan Golla, Surendra Vendra

    Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware. While high-level tensor frameworks provide flexible abstractions for model design, their eager execution models inherently lack the whole-graph visibility and granular control over hardware and memory required to maximize physical hardwar… ▽ More

    Submitted 15 July, 2026; originally announced August 2026.

  2. arXiv:2606.24780  [pdf, ps, other

    cs.AI cs.LG

    BluTrain: A C++/CUDA Framework for AI Systems

    Authors: Adhitya Charan, Adwaid Suresh, Anuj Kumar, Aparna A, Dhanakumar K, Dharun M S, Dinesh G, Goutham Kumar Reddy K, Harshini V M, Jenifa D, Jona Delcy C A, Kathirvel S, Killi Uma Maheswara Rao, Kiruthik Kanna M, Kurra Vishnu Sai, Madhumithaa G K, Navin Kumar V, Ram Charan Golla, Revathi T, Rishikkanth R, Sanjay Krishna M V, Surendra Vendra

    Abstract: Progress in deep learning is, at scale, more a matter of systems engineering than of modelling: the behaviour of a model in training (its throughput, its memory footprint, and the numerical fidelity of the result) is determined less by the architecture itself than by how that architecture is expressed on the hardware. To achieve absolute control over this hardware expression while abstracting away… ▽ More

    Submitted 6 July, 2026; v1 submitted 23 June, 2026; originally announced June 2026.