Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–9 of 9 results for author: Terekhov, M

Searching in archive cs. Search in all archives.
.
  1. arXiv:2606.08892  [pdf, ps, other

    cs.LG

    Diffuse AI Control on Fuzzy Tasks

    Authors: Mikhail Terekhov, Caglar Gulcehre, Vivek Hebbar, Joe Benton

    Abstract: AI models deployed in critical domains, such as AI safety research, may subtly sabotage our efforts due to misalignment. Diffuse AI Control is a subfield of AI safety concerned with mitigating risks from AI sabotage distributed over long deployment horizons (diffuse threats). These risks are particularly pernicious on fuzzy tasks, i.e. tasks which are hard to grade or require intuition. To underst… ▽ More

    Submitted 17 June, 2026; v1 submitted 7 June, 2026; originally announced June 2026.

  2. arXiv:2510.09462  [pdf, ps, other

    cs.LG cs.AI cs.CR

    Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols

    Authors: Mikhail Terekhov, Alexander Panfilov, Daniil Dzenhaliou, Caglar Gulcehre, Maksym Andriushchenko, Ameya Prabhu, Jonas Geiping

    Abstract: AI control protocols serve as a defense mechanism to stop untrusted LLM agents from causing harm in autonomous settings. Prior work treats this as a security problem, stress testing with exploits that use the deployment context to subtly complete harmful side tasks, such as backdoor insertion. In practice, most AI control protocols are fundamentally based on LLM monitors, which can become a centra… ▽ More

    Submitted 2 March, 2026; v1 submitted 10 October, 2025; originally announced October 2025.

  3. arXiv:2506.05296  [pdf, ps, other

    cs.AI cs.LG

    Control Tax: The Price of Keeping AI in Check

    Authors: Mikhail Terekhov, Zhen Ning David Liu, Caglar Gulcehre, Samuel Albanie

    Abstract: The rapid integration of agentic AI into high-stakes real-world applications requires robust oversight mechanisms. The emerging field of AI Control (AIC) aims to provide such an oversight mechanism, but practical adoption depends heavily on implementation overhead. To study this problem better, we introduce the notion of Control tax -- the operational and financial cost of integrating control meas… ▽ More

    Submitted 2 March, 2026; v1 submitted 5 June, 2025; originally announced June 2025.

  4. arXiv:2503.11703  [pdf, ps, other

    cs.LG

    Physical knowledge improves prediction of EM Fields

    Authors: Andrzej Dulny, Farzad Jabbarigargari, Andreas Hotho, Laura Maria Schreiber, Maxim Terekhov, Anna Krause

    Abstract: We propose a 3D U-Net model to predict the spatial distribution of electromagnetic fields inside a radio-frequency (RF) coil with a subject present, using the phase, amplitude, and position of the coils, along with the density, permittivity, and conductivity of the surrounding medium as inputs. To improve accuracy, we introduce a physics-augmented variant, U-Net Phys, which incorporates Gauss's la… ▽ More

    Submitted 12 March, 2025; originally announced March 2025.

  5. arXiv:2502.02619  [pdf, other

    q-fin.PM cs.LG q-fin.RM

    Regret-Optimized Portfolio Enhancement through Deep Reinforcement Learning and Future Looking Rewards

    Authors: Daniil Karzanov, Rubén Garzón, Mikhail Terekhov, Caglar Gulcehre, Thomas Raffinot, Marcin Detyniecki

    Abstract: This paper introduces a novel agent-based approach for enhancing existing portfolio strategies using Proximal Policy Optimization (PPO). Rather than focusing solely on traditional portfolio construction, our approach aims to improve an already high-performing strategy through dynamic rebalancing driven by PPO and Oracle agents. Our target is to enhance the traditional 60/40 benchmark (60% stocks,… ▽ More

    Submitted 4 February, 2025; originally announced February 2025.

    Comments: 11 pages, 7 figures

  6. arXiv:2501.06258  [pdf, other

    cs.LG

    Contextual Bandit Optimization with Pre-Trained Neural Networks

    Authors: Mikhail Terekhov

    Abstract: Bandit optimization is a difficult problem, especially if the reward model is high-dimensional. When rewards are modeled by neural networks, sublinear regret has only been shown under strong assumptions, usually when the network is extremely wide. In this thesis, we investigate how pre-training can help us in the regime of smaller models. We consider a stochastic contextual bandit with the rewards… ▽ More

    Submitted 9 January, 2025; originally announced January 2025.

    Comments: Master's thesis

  7. arXiv:2410.22366  [pdf, ps, other

    cs.LG cs.AI cs.CV

    One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models

    Authors: Viacheslav Surkov, Chris Wendler, Antonio Mari, Mikhail Terekhov, Justin Deschenaux, Robert West, Caglar Gulcehre, David Bau

    Abstract: For large language models (LLMs), sparse autoencoders (SAEs) have been shown to decompose intermediate representations that often are not interpretable directly into sparse sums of interpretable features, facilitating better control and subsequent analysis. However, similar analyses and approaches have been lacking for text-to-image models. We investigate the possibility of using SAEs to learn int… ▽ More

    Submitted 27 October, 2025; v1 submitted 28 October, 2024; originally announced October 2024.

  8. arXiv:2407.16807  [pdf, other

    cs.LG cs.AI

    In Search for Architectures and Loss Functions in Multi-Objective Reinforcement Learning

    Authors: Mikhail Terekhov, Caglar Gulcehre

    Abstract: Multi-objective reinforcement learning (MORL) is essential for addressing the intricacies of real-world RL problems, which often require trade-offs between multiple utility functions. However, MORL is challenging due to unstable learning dynamics with deep learning-based function approximators. The research path most taken has been to explore different value-based loss functions for MORL to overco… ▽ More

    Submitted 23 July, 2024; originally announced July 2024.

    Comments: 20 pages, 10 figures, 3 tables

  9. arXiv:2406.18370  [pdf, ps, other

    quant-ph cs.AI cs.LG stat.ML

    Learning pure quantum states (almost) without regret

    Authors: Josep Lumbreras, Mikhail Terekhov, Marco Tomamichel

    Abstract: We initiate the study of sample-optimal quantum state tomography with minimal disturbance to the samples. Can we efficiently learn a precise description of a quantum state through sequential measurements of samples while at the same time making sure that the post-measurement state of the samples is only minimally perturbed? Defining regret as the cumulative disturbance of all samples, the challeng… ▽ More

    Submitted 5 June, 2025; v1 submitted 26 June, 2024; originally announced June 2024.

    Comments: 28 pages, 2 figures