Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 209 results for author: Santos, C

Searching in archive cs. Search in all archives.
.
  1. arXiv:2607.22890  [pdf, ps, other

    cs.GR cs.CV cs.LG

    Meshless Domain Randomization via Explicit Parameter Perturbation of 3D Gaussian Splatting

    Authors: Felipe Nunes Carbone de Carvalho, Joyce de Morais Souza, Alan de Aguiar, Charles Morphy D. Santos, João Paulo Gois

    Abstract: Domain Randomization (DR) is a standard technique for closing the Sim-to-Real gap, yet traditional DR pipelines rely on classical computer graphics rendering driven by polygon meshes. For complex organic subjects, such as insect specimens, extracting and rendering textured meshes is challenging. To address this issue, we propose a meshless DR framework that operates on the parameter space of 3D Ga… ▽ More

    Submitted 24 July, 2026; originally announced July 2026.

    Comments: 10 pages, 4 figures

  2. arXiv:2607.21904  [pdf, ps, other

    cs.CV cs.CL

    Diffusion Models in Medical Image Inpainting: Challenges, Solution Taxonomy, and Future Directions

    Authors: Arthur Dantas Mangussi, Joana Cristo Santos, Ricardo Cardoso Pereira, Ana Carolina Lorena, Mário A. T. Figueiredo, Pedro Henriques Abreu

    Abstract: Image inpainting aims to reconstruct missing or corrupted regions of an image while preserving as much as possible, visual and semantic consistency. In medical imaging, this task is particularly important because artifacts, missing information, and pathological alterations can compromise diagnostic reliability and downstream clinical applications. Recently, diffusion models have emerged as state-o… ▽ More

    Submitted 23 July, 2026; originally announced July 2026.

  3. arXiv:2607.14124  [pdf, ps, other

    stat.AP cs.LG

    Analysis of Public Schools Educational Performance Based on Causal Models and Hierarchical Clustering

    Authors: Anderson L. de Paula, Pedro C. dos Santos, Renato A. Krohling

    Abstract: The increasing availability of large-scale educational datasets has expanded the use of quantitative methods for investigating school performance. However, institutional heterogeneity among schools and the structural complexity of educational data pose substantial challenges to traditional statistical modeling approaches. This study investigates the existence of school typologies based on structur… ▽ More

    Submitted 18 June, 2026; originally announced July 2026.

    Comments: 15 pages

  4. arXiv:2606.31159  [pdf, ps, other

    cs.SE cs.CR cs.LG

    An Empirical Study of Security Calibration in Large Language Models for Code

    Authors: Mohammed Latif Siddiq, Md. Nafiu Rahman, Joanna C. S. Santos

    Abstract: Large Language Models (LLMs) are rapidly transforming software development, yet their use in security-critical contexts raises a key question: do models know when their generated code is insecure? This property, known as calibration, measures whether a model's confidence aligns with the true correctness of its outputs. We present the first large-scale empirical study of security calibration in LLM… ▽ More

    Submitted 30 June, 2026; originally announced June 2026.

    Comments: Accepted at the 42nd International Conference on Software Maintenance and Evolution (ICSME 2026) Research Track

  5. arXiv:2606.25863  [pdf

    cs.SE cs.CR

    Automated Detection of Configuration-Specific Security Vulnerabilities via Patch Analysis

    Authors: Felipe de Sant'Anna Paixão, Joanna C. S. Santos, Paulo Anselmo da Mota Silveira Neto, Daniel Sadoc Menasche, Gustavo Bittencourt Figueiredo, Eduardo Santana de Almeida

    Abstract: We study how security patches in highly configurable C/C++ systems map onto the space of compile-time variants. We formalize the Vulnerability Impact Condition (VIC) - a Boolean predicate over configuration options that denotes all variants that contained the original flaw - and introduce PatchLens, a purely static technique that recovers VICs by aligning AST-level patch hunks with source-level pr… ▽ More

    Submitted 24 June, 2026; originally announced June 2026.

    Comments: Accepted at FSE 2026

  6. arXiv:2605.18634  [pdf, ps, other

    math.CO cs.DM

    Harmonious Colorings: bounds, heuristics and integer-linear formulations

    Authors: Júlio Araújo, Manoel Campêlo, Beatriz Martins, Marcio C. Santos

    Abstract: A proper coloring $c$ of a simple graph $G$ is harmonious if, for every pair of distinct edges $uv,xy\in E(G)$, we have that $\{c(u),c(v)\}\neq \{c(x),c(y)\}$. The harmonious chromatic number of $G$, denoted by $h(G)$, is the least positive integer $k$ such that $G$ has a harmonious coloring with $k$ colors. In this work, we extend an idea presented in [Kolay, et al. Harmonious coloring: Parameter… ▽ More

    Submitted 18 May, 2026; originally announced May 2026.

    Comments: 20 pages, 1 figure, 6 tables

    MSC Class: 68R10

  7. OrganicHAR: Towards Activity Discovery in Organic Settings for Privacy Preserving Sensors Using Efficient Video Analysis

    Authors: Prasoon Patidar, Riku Arakawa, Ricardo Graça, Rúben Moutinho, Adriano Soares, Ana Vasconcelos, Filippo Talami, Joana Couto da Silva, Inês Silva, Cristina Mendes Santos, Mayank Goel, Yuvraj Agarwal

    Abstract: Deploying human activity recognition (HAR) at home is still rare because sensor signals vary wildly across houses, people, and time, essentially requiring in-situ data collection and training. Prior approaches use cameras to generate training labels for privacy-preserving sensors (LiDAR, RADAR, Thermal), but this forces sensors to detect predefined activities that cameras can see yet the sensors t… ▽ More

    Submitted 18 May, 2026; originally announced May 2026.

    Comments: 23 pages, 14 figures, with a 4-page appendix containing 2 additional figures. To appear in Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. (IMWUT), Vol. 9, No. 4, Article 203 (December 2025). DOI: 10.1145/3770674

    ACM Class: H.5.2; I.2.10; I.5.4

    Journal ref: Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 9(4) (2025) 203:1-203:32

  8. arXiv:2604.10358  [pdf, ps, other

    cs.RO

    COSMIK-MPPI: Scaling Constrained Model Predictive Control to Collision Avoidance in Close-Proximity Dynamic Human Environments

    Authors: Ege Gursoy, Maxime Sabbah, Arthur Haffemayer, Joao Cavalcanti Santos, Pietro Noah Crestaz, Vladimir Petrik, Nicolas Mansard, Vincent Bonnet

    Abstract: Ensuring safe physical interaction between torque-controlled manipulators and humans is essential for deploying robots in everyday environments. Model Predictive Control (MPC) has emerged as a suitable framework thanks to its capacity to handle hard constraints, provide strong guarantees and zero-shot adaptability through predictive reasoning. However, Gradient-Based MPC (GB-MPC) solvers have demo… ▽ More

    Submitted 11 April, 2026; originally announced April 2026.

  9. arXiv:2604.00607  [pdf, ps, other

    quant-ph cs.DS math.OC

    Beyond Quantum Advantage: Improved Classical Algorithms for the Binary Paint Shop Problem

    Authors: Mark Goh, Lara Caroline Pereira dos Santos, Thorge Müller, Matthias Sperl

    Abstract: The binary paint shop problem (BPSP) is an APX-hard optimization problem in which, given $n$ car models that occur twice in a sequence of length $2n$, the objective is to find a colouring sequence such that each car model pair is painted differently while minimizing the number of times the paint is swapped along the sequence. A recent classical heuristic, known as the recursive star greedy (RSG) a… ▽ More

    Submitted 19 August, 2026; v1 submitted 1 April, 2026; originally announced April 2026.

    Comments: 11 pages, 4 figures. Additional results for V3 including improved exact upper-bound rather than numerical upper-bound

  10. arXiv:2603.30040  [pdf, ps, other

    cs.SE cs.AI

    Automatic Identification of Parallelizable Loops Using Transformer-Based Source Code Representations

    Authors: Izavan dos S. Correia, Henrique C. T. Santos, Tiago A. E. Ferreira

    Abstract: Automatic parallelization remains a challenging problem in software engineering, particularly in identifying code regions where loops can be safely executed in parallel on modern multi-core architectures. Traditional static analysis techniques, such as dependence analysis and polyhedral models, often struggle with irregular or dynamically structured code. In this work, we propose a Transformer-bas… ▽ More

    Submitted 31 March, 2026; originally announced March 2026.

    Comments: 28 pages, 12 figures

  11. arXiv:2603.13636  [pdf, ps, other

    cs.CL cs.AI cs.CY

    Widespread Gender and Pronoun Bias in Moral Judgments Across LLMs

    Authors: Gustavo Lúcius Fernandes, Jeiverson C. V. M. Santos, Pedro O. S. Vaz-de-Melo

    Abstract: Large language models (LLMs) are increasingly used to assess moral or ethical statements, yet their judgments may reflect social and linguistic biases. This work presents a controlled, sentence-level study of how grammatical person, number, and gender markers influence LLM moral classifications of fairness. Starting from 550 balanced base sentences from the ETHICS dataset, we generated 26 counterf… ▽ More

    Submitted 13 March, 2026; originally announced March 2026.

  12. arXiv:2603.00079  [pdf, ps, other

    cs.CY

    Dynamic Map-based Data-Centric Approach for Tourism and Cultural Heritage Preservation Digital Twins

    Authors: João Spínola Falcão, João G. Perrone Hohlenwerger, Daniel C. Santos, Lucas Almeida de Sousa, Nazim Agoulmine, Joberto S. B. Martins

    Abstract: Tourism is an essential and growing economic activity worldwide, bringing benefits such as job creation, revenue generation, and tax revenue, and driving economic prosperity. Tourism activities may also have negative impacts on cities, including overtourism and pressure on housing and real estate markets. Cultural heritage is an essential asset of cities and countries that must be preserved. Cultu… ▽ More

    Submitted 13 February, 2026; originally announced March 2026.

    Comments: 8 pages, 5 figures, ADVANCE 26 - International Workshop on ADVANCEs in ICT Infrastructures and Services - March 25th-27th, 2026

  13. arXiv:2602.20222  [pdf, ps, other

    cs.CR cs.CY

    The TCF doesn't really A(A)ID -- Automatic Privacy Analysis and Legal Compliance of TCF-based Android Applications

    Authors: Victor Morel, Cristiana Santos, Pontus Carlsson, Joel Ahlinder, Romaric Duvignau

    Abstract: The Transparency and Consent Framework (TCF), developed by the Interactive Advertising Bureau (IAB) Europe, provides a de facto standard for requesting, recording, and managing user consent from European end-users. This framework has previously been found to infringe European data protection law and has subsequently been regularly updated. Previous research on the TCF focused exclusively on web co… ▽ More

    Submitted 27 February, 2026; v1 submitted 23 February, 2026; originally announced February 2026.

    Comments: Accepted for publication at PETS'26

  14. Characterizing and Modeling the GitHub Security Advisories Review Pipeline

    Authors: Claudio Segal, Paulo Segal, Carlos Eduardo Banjar, Felipe de Sant'Anna Paixão, Hudson Silva Borges, Paulo Silveira, Eduardo Santana de Almeida, Joanna C. S. Santos, Anton Kocheturov, Gaurav Kumar Srivastava, Daniel Sadoc Menasché

    Abstract: GitHub Security Advisories (GHSA) have become a central component of open-source vulnerability disclosure and are widely used by developers and security tools. A distinctive feature of GHSA is that only a fraction of advisories are reviewed by GitHub, while the mechanisms associated with this review process remain poorly understood. In this paper, we conduct a large-scale empirical study of the GH… ▽ More

    Submitted 30 April, 2026; v1 submitted 5 February, 2026; originally announced February 2026.

    Comments: Paper accepted at 23rd International Mining Software Repositories Conference (MSR 2026)

  15. arXiv:2602.02509  [pdf, ps, other

    cs.CY cs.AI

    CodeGuard: Improving LLM Guardrails in CS Education

    Authors: Nishat Raihan, Noah Erdachew, Jayoti Devi, Joanna C. S. Santos, Marcos Zampieri

    Abstract: Large language models (LLMs) are increasingly embedded in Computer Science (CS) classrooms to automate code generation, feedback, and assessment. However, their susceptibility to adversarial or ill-intentioned prompts threatens student learning and academic integrity. To cope with this important issue, we evaluate existing off-the-shelf LLMs in handling unsafe and irrelevant prompts within the dom… ▽ More

    Submitted 22 January, 2026; originally announced February 2026.

  16. arXiv:2601.14163  [pdf, ps, other

    cs.SE cs.CR

    An Empirical Study on Remote Code Execution in Machine Learning Model Hosting Ecosystems

    Authors: Mohammed Latif Siddiq, Tanzim Hossain Romel, Natalie Sekerak, Beatrice Casey, Joanna C. S. Santos

    Abstract: Model-sharing platforms, such as Hugging Face, ModelScope, and OpenCSG, have become central to modern machine learning development, enabling developers to share, load, and fine-tune pre-trained models with minimal effort. However, the flexibility of these ecosystems introduces a critical security concern: the execution of untrusted code during model loading (i.e., via trust_remote_code or trust_re… ▽ More

    Submitted 20 January, 2026; originally announced January 2026.

  17. arXiv:2601.00477  [pdf, ps, other

    cs.CR cs.SE

    Security in the Age of AI Teammates: An Empirical Study of Agentic Pull Requests on GitHub

    Authors: Mohammed Latif Siddiq, Xinye Zhao, Vinicius Carvalho Lopes, Beatrice Casey, Joanna C. S. Santos

    Abstract: Autonomous coding agents are increasingly deployed as AI teammates in modern software engineering, independently authoring pull requests (PRs) that modify production code at scale. This study aims to systematically characterize how autonomous coding agents contribute to software security in practice, how these security-related contributions are reviewed and accepted, and which observable signals a… ▽ More

    Submitted 1 January, 2026; originally announced January 2026.

    Comments: Submitted to Information and Software Technology Journal

  18. arXiv:2512.21238  [pdf, ps, other

    cs.SE cs.CR cs.LG

    Assessing the Software Security Comprehension of Large Language Models

    Authors: Mohammed Latif Siddiq, Natalie Sekerak, Antonio Karam, Maria Leal, Arvin Islam-Gomes, Joanna C. S. Santos

    Abstract: Large language models (LLMs) are increasingly used in software development, but their level of software security expertise remains unclear. This work systematically evaluates the security comprehension of five leading LLMs: GPT-4o-Mini, GPT-5-Mini, Gemini-2.5-Flash, Llama-3.1, and Qwen-2.5, using Blooms Taxonomy as a framework. We assess six cognitive dimensions: remembering, understanding, applyi… ▽ More

    Submitted 24 December, 2025; originally announced December 2025.

    Comments: Submitted to Empirical Software Engineering (EMSE) journal

  19. arXiv:2512.13910  [pdf, ps, other

    cs.LG cs.AI

    Exploring Machine Learning, Deep Learning, and Explainable AI Methods for Seasonal Precipitation Prediction in South America

    Authors: Matheus Corrêa Domingos, Valdivino Alexandre de Santiago Júnior, Juliana Aparecida Anochi, Elcio Hideiti Shiguemori, Luísa Mirelle Costa dos Santos, Hércules Carlos dos Santos Pereira, André Estevam Costa Oliveira

    Abstract: Forecasting meteorological variables is challenging due to the complexity of their processes, requiring advanced models for accuracy. Accurate precipitation forecasts are vital for society. Reliable predictions help communities mitigate climatic impacts. Based on the current relevance of artificial intelligence (AI), classical machine learning (ML) and deep learning (DL) techniques have been used… ▽ More

    Submitted 15 December, 2025; originally announced December 2025.

  20. Can the GPC standard eliminate consent banners in the EU?

    Authors: Sebastian Zimmeck, Harshvardhan J. Pandit, Frederik Zuiderveen Borgesius, Cristiana Teixeira Santos, Konrad Kollnig, Robin Berjon

    Abstract: In the EU, the General Data Protection Regulation and the ePrivacy Directive mandate consent for the use of personal data for the purpose of behavioural advertising and tracking technologies. However, the ubiquity of consent banners has led to widespread consent fatigue and questions about the effectiveness of these mechanisms in protecting data subjects' data. To simplify digital laws and make th… ▽ More

    Submitted 6 May, 2026; v1 submitted 9 December, 2025; originally announced December 2025.

    Journal ref: Computer Law & Security Review, 61, 106332 (2026)

  21. arXiv:2512.00651  [pdf, ps, other

    cs.SE cs.LG

    Large Language Models for Software Engineering: A Reproducibility Crisis

    Authors: Mohammed Latif Siddiq, Arvin Islam-Gomes, Natalie Sekerak, Joanna C. S. Santos

    Abstract: Reproducibility is a cornerstone of scientific progress, yet its state in large language model (LLM)-based software engineering (SE) research remains poorly understood. This paper presents the first large-scale, empirical study of reproducibility practices in LLM-for-SE research. We systematically mined and analyzed 640 papers published between 2017 and 2025 across premier software engineering, ma… ▽ More

    Submitted 29 November, 2025; originally announced December 2025.

    Comments: Submitted to Empirical Software Engineering (EMSE) journal; 112 pages (81 pages of references)

  22. arXiv:2511.12822  [pdf, ps, other

    cs.CY

    The Unspoken Crisis of Learning: The Surging Zone of No Development

    Authors: Euzeli C. dos Santos Jr., Tracey Birdwell

    Abstract: AI has redefined the boundaries of assistance in education, often blurring the line between guided learning and dependency. This paper revisits Vygotsky's Zone of Proximal Development (ZPD) through the lens of the P2P Teaching framework. By contrasting temporary scaffolding with the emerging phenomenon of permanent digital mediation, the study introduces the concept of the Zone of No Development (… ▽ More

    Submitted 16 November, 2025; originally announced November 2025.

    Comments: 6 pages, 3 figures

  23. arXiv:2511.08059  [pdf, ps, other

    cs.SE cs.CR cs.HC

    "I need to learn better searching tactics for privacy policy laws." Investigating Software Developers' Behavior When Using Sources on Privacy Issues

    Authors: Stefan Albert Horstmann, Sandy Hong, Maziar Niazian, Cristiana Santos, Alena Naiakshina

    Abstract: Since the introduction of the European General Data Protection Regulation (GDPR) and the California Consumer Privacy Act (CCPA), software developers increasingly have to make privacy-related decisions during system design and implementation. However, past research showed that they often lack legal expertise and struggle with privacy-compliant development. To shed light on how effective current inf… ▽ More

    Submitted 8 December, 2025; v1 submitted 11 November, 2025; originally announced November 2025.

    Journal ref: 2026 IEEE/ACM 48th International Conference on Software Engineering (ICSE26)

  24. arXiv:2511.04426  [pdf, ps, other

    cs.CV

    HideAndSeg: an AI-based tool with automated prompting for octopus segmentation in natural habitats

    Authors: Alan de Aguiar, Michaella Pereira Andrade, Charles Morphy D. Santos, João Paulo Gois

    Abstract: Analyzing octopuses in their natural habitats is challenging due to their camouflage capability, rapid changes in skin texture and color, non-rigid body deformations, and frequent occlusions, all of which are compounded by variable underwater lighting and turbidity. Addressing the lack of large-scale annotated datasets, this paper introduces HideAndSeg, a novel, minimally supervised AI-based tool… ▽ More

    Submitted 6 November, 2025; originally announced November 2025.

  25. Investigating Software Aging in LLM-Generated Software Systems

    Authors: César Santos, Ermeson Andrade, Roberto Natella

    Abstract: Automatically generated software, especially code produced by Large Language Models (LLMs), is increasingly adopted to accelerate development and reduce manual effort. However, little is known about the long-term reliability of such systems under sustained execution. In this paper, we experimentally investigate the phenomenon of software aging in applications generated by LLM-based tools. Using th… ▽ More

    Submitted 28 October, 2025; originally announced October 2025.

    Comments: Presented at the 17th International Workshop on Software Aging and Rejuvenation (WoSAR), 2025

  26. arXiv:2510.23817  [pdf, ps, other

    cs.LG stat.ME

    Combining SHAP and Causal Analysis for Interpretable Fault Detection in Industrial Processes

    Authors: Pedro Cortes dos Santos, Matheus Becali Rocha, Renato A Krohling

    Abstract: Industrial processes generate complex data that challenge fault detection systems, often yielding opaque or underwhelming results despite advanced machine learning techniques. This study tackles such difficulties using the Tennessee Eastman Process, a well-established benchmark known for its intricate dynamics, to develop an innovative fault detection framework. Initial attempts with standard mode… ▽ More

    Submitted 27 October, 2025; originally announced October 2025.

  27. arXiv:2509.16215  [pdf, ps, other

    cs.LG cs.AI cs.DC cs.NE cs.PL cs.SE

    Discovering Software Parallelization Points Using Deep Neural Networks

    Authors: Izavan dos S. Correia, Henrique C. T. Santos, Tiago A. E. Ferreira

    Abstract: This study proposes a deep learning-based approach for discovering loops in programming code according to their potential for parallelization. Two genetic algorithm-based code generators were developed to produce two distinct types of code: (i) independent loops, which are parallelizable, and (ii) ambiguous loops, whose dependencies are unclear, making them impossible to define if the loop is para… ▽ More

    Submitted 1 October, 2025; v1 submitted 5 September, 2025; originally announced September 2025.

    Comments: 17 pages, 10 figures

  28. arXiv:2509.06452  [pdf, ps, other

    cs.IR cs.SD

    AudioBoost: Increasing Audiobook Retrievability in Spotify Search with Synthetic Query Generation

    Authors: Enrico Palumbo, Gustavo Penha, Alva Liu, Marcus Eltscheminov, Jefferson Carvalho dos Santos, Alice Wang, Hugues Bouchard, Humberto Jesús Corona Pampin, Michelle Tran Luu

    Abstract: Spotify has recently introduced audiobooks as part of its catalog, complementing its music and podcast offering. Search is often the first entry point for users to access new items, and an important goal for Spotify is to support users in the exploration of the audiobook catalog. More specifically, we would like to enable users without a specific item in mind to broadly search by topic, genre, sto… ▽ More

    Submitted 8 September, 2025; originally announced September 2025.

    Comments: EARL Workshop @ RecSys25

  29. arXiv:2508.12520  [pdf, ps, other

    cs.CV cs.AI

    An Initial Study of Bird's-Eye View Generation for Autonomous Vehicles using Cross-View Transformers

    Authors: Felipe Carlos dos Santos, Eric Aislan Antonelo, Gustavo Claudio Karl Couto

    Abstract: Bird's-Eye View (BEV) maps provide a structured, top-down abstraction that is crucial for autonomous-driving perception. In this work, we employ Cross-View Transformers (CVT) for learning to map camera images to three BEV's channels - road, lane markings, and planned trajectory - using a realistic simulator for urban driving. Our study examines generalization to unseen towns, the effect of differe… ▽ More

    Submitted 17 August, 2025; originally announced August 2025.

    Comments: 12 pages,submitted in ENIAC 2025

  30. arXiv:2507.16164  [pdf, ps, other

    cs.CR cs.AI cs.LG

    Attacking interpretable NLP systems

    Authors: Eldor Abdukhamidov, Tamer Abuhmed, Joanna C. S. Santos, Mohammed Abuhamad

    Abstract: Studies have shown that machine learning systems are vulnerable to adversarial examples in theory and practice. Where previous attacks have focused mainly on visual models that exploit the difference between human and machine perception, text-based models have also fallen victim to these attacks. However, these attacks often fail to maintain the semantic meaning of the text and similarity. This pa… ▽ More

    Submitted 21 July, 2025; originally announced July 2025.

    ACM Class: I.2.7; I.2.6; I.2.3; D.4.6

  31. arXiv:2507.06261  [pdf, ps, other

    cs.CL cs.AI

    Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

    Authors: Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, Luke Marris, Sam Petulla, Colin Gaffney, Asaf Aharoni, Nathan Lintz, Tiago Cardal Pais, Henrik Jacobsson, Idan Szpektor, Nan-Jiang Jiang, Krishna Haridasan, Ahmed Omran, Nikunj Saunshi, Dara Bahri, Gaurav Mishra, Eric Chu , et al. (3410 additional authors not shown)

    Abstract: In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our most capable model yet, achieving SoTA performance on frontier coding and reasoning benchmarks. In addition to its incredible coding and reasoning skills, Gemini 2.5 Pro is a thinking model that excels at multimodal unde… ▽ More

    Submitted 19 December, 2025; v1 submitted 7 July, 2025; originally announced July 2025.

    Comments: 72 pages, 17 figures

  32. arXiv:2506.06850  [pdf, ps, other

    cs.CV eess.SP

    Deep Inertial Pose: A deep learning approach for human pose estimation

    Authors: Sara M. Cerqueira, Manuel Palermo, Cristina P. Santos

    Abstract: Inertial-based Motion capture system has been attracting growing attention due to its wearability and unsconstrained use. However, accurate human joint estimation demands several complex and expertise demanding steps, which leads to expensive software such as the state-of-the-art MVN Awinda from Xsens Technologies. This work aims to study the use of Neural Networks to abstract the complex biomecha… ▽ More

    Submitted 7 June, 2025; originally announced June 2025.

  33. arXiv:2506.04260  [pdf, ps, other

    cs.CY cs.HC

    Turning to Online Forums for Legal Information: A Case Study of GDPR's Legitimate Interests

    Authors: Lin Kyi, Cristiana Santos, Sushil Ammanaghatta Shivakumar, Franziska Roesner, Asia Biega

    Abstract: Practitioners building online services and tools often turn to online forums such as Reddit, Law Stack Exchange, and Stack Overflow for legal guidance to ensure compliance with the GDPR. The legal information presented in these forums directly impacts present-day industry practitioner's decisions. Online forums can serve as gateways that, depending on the accuracy and quality of the answers provid… ▽ More

    Submitted 23 July, 2025; v1 submitted 2 June, 2025; originally announced June 2025.

    Comments: Accepted at Annual Privacy Forum 2025

  34. arXiv:2505.14700  [pdf, ps, other

    cs.LG math.NA stat.ML

    Stochastic Fractional Neural Operators: A Symmetrized Approach to Modeling Turbulence in Complex Fluid Dynamics

    Authors: Rômulo Damasclin Chaves dos Santos, Jorge Henrique de Oliveira Sales

    Abstract: In this work, we introduce a new class of neural network operators designed to handle problems where memory effects and randomness play a central role. In this work, we introduce a new class of neural network operators designed to handle problems where memory effects and randomness play a central role. These operators merge symmetrized activation functions, Caputo-type fractional derivatives, and… ▽ More

    Submitted 12 May, 2025; originally announced May 2025.

    Comments: 17 pages

  35. arXiv:2505.12892  [pdf, ps, other

    cs.CY

    "I will never pay for this" Perception of fairness and factors affecting behaviour on 'pay-or-ok' models

    Authors: Victor Morel, Farzaneh Karegar, Cristiana Santos

    Abstract: The rise of cookie paywalls ('pay-or-ok' models) has prompted growing debates around the right to privacy and data protection, monetisation, and the legitimacy of user consent. Despite their increasing use across sectors, limited research has explored how users perceive these models or what shapes their decisions to either consent to tracking or pay. To address this gap, we conducted four focus gr… ▽ More

    Submitted 31 July, 2025; v1 submitted 19 May, 2025; originally announced May 2025.

    Comments: Accepted for publication at APF2025

  36. arXiv:2504.11406  [pdf, other

    cs.CV cs.AI

    Multi-level Cellular Automata for FLIM networks

    Authors: Felipe Crispim Salvagnini, Jancarlo F. Gomes, Cid A. N. Santos, Silvio Jamil F. Guimarães, Alexandre X. Falcão

    Abstract: The necessity of abundant annotated data and complex network architectures presents a significant challenge in deep-learning Salient Object Detection (deep SOD) and across the broader deep-learning landscape. This challenge is particularly acute in medical applications in developing countries with limited computational resources. Combining modern and classical techniques offers a path to maintaini… ▽ More

    Submitted 15 April, 2025; originally announced April 2025.

  37. arXiv:2504.10294  [pdf, other

    cs.RO

    Ankle Exoskeletons in Walking and Load-Carrying Tasks: Insights into Biomechanics and Human-Robot Interaction

    Authors: J. F. Almeida, J. André, C. P. Santos

    Abstract: Background: Lower limb exoskeletons can enhance quality of life, but widespread adoption is limited by the lack of frameworks to assess their biomechanical and human-robot interaction effects, which are essential for developing adaptive and personalized control strategies. Understanding impacts on kinematics, muscle activity, and HRI dynamics is key to achieve improved usability of wearable robots… ▽ More

    Submitted 14 April, 2025; originally announced April 2025.

  38. arXiv:2504.10102  [pdf, ps, other

    cs.RO eess.SY

    A Human-Sensitive Controller: Adapting to Human Musculoskeletal Disorder-Related Constraints via Reinforcement Learning

    Authors: Vitor Martins, Sara M. Cerqueira, Mercedes Balcells, Elazer R Edelman, Cristina P. Santos

    Abstract: Work-Related Musculoskeletal Disorders continue to be a major challenge in industrial environments, leading to reduced workforce participation, increased healthcare costs, and long-term disability. This study introduces a human-sensitive robotic system aimed at reintegrating individuals with a history of musculoskeletal disorders into standard job roles, while simultaneously optimizing ergonomic c… ▽ More

    Submitted 5 June, 2026; v1 submitted 14 April, 2025; originally announced April 2025.

  39. arXiv:2504.03751  [pdf, ps, other

    cs.LG

    Revolutionizing Fractional Calculus with Neural Networks: Voronovskaya-Damasclin Theory for Next-Generation AI Systems

    Authors: Rômulo Damasclin Chaves dos Santos, Jorge Henrique de Oliveira Sales

    Abstract: This work introduces rigorous convergence rates for neural network operators activated by symmetrized and perturbed hyperbolic tangent functions, utilizing novel Voronovskaya-Damasclin asymptotic expansions. We analyze basic, Kantorovich, and quadrature-type operators over infinite domains, extending classical approximation theory to fractional calculus via Caputo derivatives. Key innovations incl… ▽ More

    Submitted 1 April, 2025; originally announced April 2025.

    Comments: 16 pages

  40. Impacto de Treinamento em Programação Competitiva no Ensino Médio: Resultados e Desafios

    Authors: Camila da Cruz Santos, Sarah Souto dos Santos, Crishna Irion, Giullia Rodrigues de Menezes, Rafael Dias Araújo, João Henrique de Souza Pereira

    Abstract: This article presents an ongoing research aiming to develop an effective methodology for teaching programming, focusing on participation in the Brazilian Informatics Olympiad (OBI), for elementary and high school students. The training conducted with students from the Federal Institute and state schools, demonstrates the importance of programming training programs as a way to promote interest in c… ▽ More

    Submitted 31 January, 2025; originally announced March 2025.

    Comments: 10 pages, in Portuguese, 8 figures and 2 tables

    Journal ref: Simpósio Brasileiro de Informática na Educação (SBIE) 2024

  41. Promoting Gender Equality in Competitive Programming: Strategies and Impacts of Affirmative Actions in Programming Marathons in Brazil

    Authors: Crishna Irion, Camila da Cruz Santos, Luiz Claudio Theodoro, Rafael Dias Araujo, Joao Henrique de Souza Pereira

    Abstract: In the context of Computing, competitive programming is a relevant area that aims to have students, usually in teams, solve programming challenges, developing skills and competencies in the field. However, female participation remains significantly low and notably distant compared to male participation, even with proven intellectual equity between genders. This research aims to present strategies… ▽ More

    Submitted 21 February, 2025; originally announced February 2025.

    Comments: 12 pages, SBIE (2024), in Portuguese language

  42. Humanity's Last Exam

    Authors: Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, Michael Choi, Anish Agrawal, Arnav Chopra, Adam Khoja, Ryan Kim, Richard Ren, Jason Hausenloy, Oliver Zhang, Mantas Mazeika, Dmitry Dodonov, Tung Nguyen, Jaeho Lee, Daron Anderson, Mikhail Doroshenko, Alun Cennyth Stokes , et al. (1133 additional authors not shown)

    Abstract: Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achieve over 90\% accuracy on popular benchmarks like MMLU, limiting informed measurement of state-of-the-art LLM capabilities. In response, we introduce Humanity's Last Exam (HLE), a multi-modal benchmark at the frontier of… ▽ More

    Submitted 28 July, 2026; v1 submitted 24 January, 2025; originally announced January 2025.

    Comments: 29 pages, 6 figures

  43. arXiv:2501.10496  [pdf, ps, other

    stat.ML cs.LG

    Extension of Symmetrized Neural Network Operators with Fractional and Mixed Activation Functions

    Authors: Rômulo Damasclin Chaves dos Santos, Jorge Henrique de Oliveira Sales

    Abstract: We propose a novel extension to symmetrized neural network operators by incorporating fractional and mixed activation functions. This study addresses the limitations of existing models in approximating higher-order smooth functions, particularly in complex and high-dimensional spaces. Our framework introduces a fractional exponent in the activation functions, allowing adaptive non-linear approxima… ▽ More

    Submitted 17 January, 2025; originally announced January 2025.

    Comments: 13 pages

  44. arXiv:2501.02170  [pdf

    cs.SE

    An Empirical Study of Safetensors' Usage Trends and Developers' Perceptions

    Authors: Beatrice Casey, Kaia Damian, Andrew Cotaj, Joanna C. S. Santos

    Abstract: Developers are sharing pre-trained Machine Learning (ML) models through a variety of model sharing platforms, such as Hugging Face, in an effort to make ML development more collaborative. To share the models, they must first be serialized. While there are many methods of serialization in Python, most of them are unsafe. To tame this insecurity, Hugging Face released safetensors as a way to mitigat… ▽ More

    Submitted 3 January, 2025; originally announced January 2025.

  45. arXiv:2411.15414  [pdf, other

    cs.CR

    Measuring Compliance of Consent Revocation on the Web

    Authors: Gayatri Priyadarsini Kancherla, Nataliia Bielova, Cristiana Santos, Abhishek Bichhawat

    Abstract: The GDPR requires websites to facilitate the right to revoke consent from Web users. While numerous studies measured compliance of consent with the various consent requirements, no prior work has studied consent revocation on the Web. Therefore, it remains unclear how difficult it is to revoke consent on the websites' interfaces, nor whether revoked consent is properly stored and communicated behi… ▽ More

    Submitted 22 May, 2025; v1 submitted 22 November, 2024; originally announced November 2024.

  46. arXiv:2411.14533  [pdf, other

    math.OC cs.DM

    The connected Grundy coloring problem: Formulations and a local-search enhanced biased random-key genetic algorithm

    Authors: Mateus C. Silva, Rafael A. Melo, Mauricio G. C. Resende, Marcio C. Santos, Rodrigo F. Toso

    Abstract: Given a graph G=(V,E), a connected Grundy coloring is a proper vertex coloring that can be obtained by a first-fit heuristic on a connected vertex sequence. A first-fit coloring heuristic is one that attributes to each vertex in a sequence the lowest-index color not used for its preceding neighbors. A connected vertex sequence is one in which each element, except for the first one, is connected to… ▽ More

    Submitted 21 November, 2024; originally announced November 2024.

    ACM Class: F.2; G.2

  47. arXiv:2410.17736  [pdf, other

    cs.CL

    MojoBench: Language Modeling and Benchmarks for Mojo

    Authors: Nishat Raihan, Joanna C. S. Santos, Marcos Zampieri

    Abstract: The recently introduced Mojo programming language (PL) by Modular, has received significant attention in the scientific community due to its claimed significant speed boost over Python. Despite advancements in code Large Language Models (LLMs) across various PLs, Mojo remains unexplored in this context. To address this gap, we introduce MojoBench, the first framework for Mojo code generation. Mojo… ▽ More

    Submitted 23 October, 2024; originally announced October 2024.

  48. arXiv:2410.16349  [pdf, other

    cs.LG cs.HC

    Large Language Models in Computer Science Education: A Systematic Literature Review

    Authors: Nishat Raihan, Mohammed Latif Siddiq, Joanna C. S. Santos, Marcos Zampieri

    Abstract: Large language models (LLMs) are becoming increasingly better at a wide range of Natural Language Processing tasks (NLP), such as text generation and understanding. Recently, these models have extended their capabilities to coding tasks, bridging the gap between natural languages (NL) and programming languages (PL). Foundational models such as the Generative Pre-trained Transformer (GPT) and LLaMA… ▽ More

    Submitted 21 October, 2024; originally announced October 2024.

    Comments: Accepted at 56th ACM Technical Symposium on Computer Science Education (SIGCSE TS 2025)

  49. arXiv:2410.16294  [pdf

    cs.CY

    Emílias Podcast -- Mulheres na Computação: Ampliando Horizontes e Inspirando Carreiras em STEM

    Authors: Nathálya Chaves Dos Santos, Adolfo Gustavo Serra Seca Neto

    Abstract: On October 3, 2024, the "Emílias Podcast -- Women in Computing" celebrates its 5th anniversary, standing out as a platform that promotes the participation of women in STEM (an acronym for "science, technology, engineering, and mathematics"). The podcast aims to provide a space for women in computing and related fields to share their experiences and highlight the various opportunities in Informatio… ▽ More

    Submitted 6 October, 2024; originally announced October 2024.

    Comments: In Portuguese. Accepted for presentation at XIV SEMINÁRIO DE EXTENSÃO E INOVAÇÃO. UTFPR 2024

  50. arXiv:2410.04490  [pdf

    cs.CR cs.LG cs.SE

    A Large-Scale Exploit Instrumentation Study of AI/ML Supply Chain Attacks in Hugging Face Models

    Authors: Beatrice Casey, Joanna C. S. Santos, Mehdi Mirakhorli

    Abstract: The development of machine learning (ML) techniques has led to ample opportunities for developers to develop and deploy their own models. Hugging Face serves as an open source platform where developers can share and download other models in an effort to make ML development more collaborative. In order for models to be shared, they first need to be serialized. Certain Python serialization methods a… ▽ More

    Submitted 6 October, 2024; originally announced October 2024.