Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 85 results for author: Sanchez, D

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.06154  [pdf, ps, other

    cs.RO cs.AI cs.CV

    Visual Grounding in Zero-Shot Vision-Language Control

    Authors: J. de Curtò, Dayani Plasencia, Diego Sánchez, I. de Zarzà

    Abstract: Vision-language models (VLMs) are increasingly used as zero-shot controllers, but successful trajectories do not necessarily show that decisions are grounded in visual input: simulator dynamics and conservative action priors can produce favourable scores without meaningful perception. We investigate this with an input-ablation battery: blind-image controls, repeated identical inputs, lane-axis ref… ▽ More

    Submitted 6 August, 2026; originally announced August 2026.

  2. arXiv:2606.14348  [pdf, ps, other

    physics.soc-ph cond-mat.stat-mech cs.CL

    Detecting Historical Turning Points in Italian Media: A Complex Systems Approach to a Diachronic News Corpus

    Authors: Dario Zarcone, Salvatore Miccichè, David Sanchez

    Abstract: The increasing availability of large-scale textual corpora has opened new possibilities for data-driven, quantitative approaches to historical analysis using Natural Language Processing (NLP). However, diachronic corpora with historical relevance from the pre-digital era remain scarce and often incomplete. We present a quantitative approach to historical analysis based on the reconstruction and ex… ▽ More

    Submitted 12 June, 2026; originally announced June 2026.

    Comments: 16 pages, 9 figures, 1 table

  3. arXiv:2605.24002  [pdf, ps, other

    physics.chem-ph cond-mat.mtrl-sci cs.AI physics.comp-ph

    Harnessing AtomisticSkills for Agentic Atomistic Research

    Authors: Bowen Deng, Bohan Li, Matthew Cox, Hoje Chun, Juno Nam, Artur Lyssenko, Sathya Edamadaka, Jurgis Ruza, Xiaochen Du, Nofit Segal, Jesus Diaz Sanchez, Mingrou Xie, Ty Perez, Yu Yao, Miguel Steiner, Sauradeep Majumdar, Charles B. Musgrave III, Anirban Chandra, Abhirup Patra, Detlef Hohl, Connor W. Coley, Ju Li, Rafael Gómez-Bombarelli

    Abstract: Computational materials science and chemistry span vast knowledge domains and fractured software ecosystems. Although large language models (LLMs) have demonstrated research capabilities, scaling monolithic agents to manage the rigor and complexity of atomistic research remains a challenge. Here, we introduce AtomisticSkills, an open-source harness framework that empowers general-purpose AI coding… ▽ More

    Submitted 18 May, 2026; originally announced May 2026.

  4. arXiv:2605.03205  [pdf, ps, other

    cond-mat.mtrl-sci cs.AI

    From Knowledge to Action: Outcomes of the 2025 Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry

    Authors: Aritra Roy, Kevin Shen, Andrew MacBride, Awwal Oladipupo, Mudassra Taskeen, Wojtek Treyde, Ruaa A. E. A. Abakar, Ahmad D. Abbas, Elsayed Abdelfatah, Abbas A. Abdullahi, Seham S. Abyah, Chahd Rahyl Adjmi, Fariha Agbere, Savyasanchi Aggarwal, Muhammad Ahmed, Tasnim Ahmed, Motasem Ajlouni, Mattias Akke, Hussein AlAdwan, Anwaar S. Alazani, Zahra A. Alharbi, Wajd A. Aljulyhi, Mohammed A. AlKubaish, Fatima A. Almahri, Sayed A. Almohri , et al. (328 additional authors not shown)

    Abstract: Large language models (LLMs) are rapidly changing how researchers in materials science and chemistry discover, organize, and act on scientific knowledge. This paper analyzes a broad set of community-developed LLM applications in an effort to identify emerging patterns in how these systems can be used across the scientific research lifecycle. We organize the projects into two complementary categori… ▽ More

    Submitted 4 May, 2026; originally announced May 2026.

    Comments: This paper reflects contributions from hundreds of researchers worldwide through an event, follow-on discussions, and project development exploring LLM applications in materials science and chemistry. While unconventional, it captures a timely, broad, and efficient community exploration of a rapidly evolving field and offers value to the arXiv community

  5. arXiv:2604.11565  [pdf, ps, other

    cs.CL cond-mat.stat-mech cs.IT physics.soc-ph

    Phonological distances for linguistic typology and the origin of Indo-European languages

    Authors: Marius Mavridis, Juan De Gregorio, Raul Toral, David Sanchez

    Abstract: We show that short-range phoneme dependencies encode large-scale patterns of linguistic relatedness, with direct implications for quantitative typology and evolutionary linguistics. Specifically, using an information-theoretic framework, we argue that phoneme sequences modeled as second-order Markov chains essentially capture the statistical correlations of a phonological system. This finding enab… ▽ More

    Submitted 13 April, 2026; originally announced April 2026.

    Comments: 27 pages, 7 figures, 2 appendices

  6. arXiv:2603.22987  [pdf, ps, other

    cs.CR cs.LG

    A Critical Review on the Effectiveness and Privacy Threats of Membership Inference Attacks

    Authors: Najeeb Jebreel, David Sánchez, Josep Domingo-Ferrer

    Abstract: Membership inference attacks (MIAs) aim to determine whether a data sample was included in a machine learning (ML) model's training set and have become the de facto standard for measuring privacy leakages in ML. We propose an evaluation framework that defines the conditions under which MIAs constitute a genuine privacy threat, and review representative MIAs against it. We find that, under the real… ▽ More

    Submitted 24 March, 2026; originally announced March 2026.

    Comments: To appear in ESORICS 2026

  7. arXiv:2603.07567  [pdf, ps, other

    cs.CR cs.LG

    Revisiting the LiRA Membership Inference Attack Under Realistic Assumptions

    Authors: Najeeb Jebreel, Mona Khalil, David Sánchez, Josep Domingo-Ferrer

    Abstract: Membership inference attacks (MIAs) have become the standard tool for evaluating privacy leakage in machine learning (ML). Among them, the Likelihood-Ratio Attack (LiRA) is widely regarded as the state of the art when sufficient shadow models are available. However, prior evaluations have often overstated the effectiveness of LiRA by attacking models overconfident on their training samples, calibr… ▽ More

    Submitted 8 March, 2026; originally announced March 2026.

    Comments: Accepted to PoPETs 2026(3)

  8. arXiv:2603.01091  [pdf, ps, other

    cs.CR

    On the Practical Feasibility of Harvest-Now, Decrypt-Later Attacks

    Authors: Javier Blanco-Romero, Florina Almenares Mendoza, Carlos García Rubio, Celeste Campo, Daniel Díaz Sánchez

    Abstract: Harvest-now, decrypt-later (HN-DL) attacks threaten today's encrypted communications by archiving ciphertext until a quantum computer can break the underlying key exchange. This paper reframes HN-DL as an economic problem, quantifying adversary costs across Transport Layer Security (TLS) 1.2, TLS 1.3, QUIC, and Secure Shell (SSH) with an open-source testbed that reproduces the full attack sequence… ▽ More

    Submitted 1 March, 2026; originally announced March 2026.

  9. arXiv:2511.16501  [pdf, ps, other

    cs.LG cs.AI

    ODE-ViT: Plug & Play Attention Layer from the Generalization of the ViT as an Ordinary Differential Equation

    Authors: Carlos Boned Riera, David Romero Sanchez, Oriol Ramos Terrades

    Abstract: In recent years, increasingly large models have achieved outstanding performance across CV tasks. However, these models demand substantial computational resources and storage, and their growing complexity limits our understanding of how they make decisions. Most of these architectures rely on the attention mechanism within Transformer-based designs. Building upon the connection between residual ne… ▽ More

    Submitted 20 November, 2025; originally announced November 2025.

  10. arXiv:2511.08702  [pdf, ps, other

    cs.LG cs.AI cs.CR cs.CY

    FAIRPLAI: A Human-in-the-Loop Approach to Fair and Private Machine Learning

    Authors: David Sanchez Jr., Holly Lopez, Michelle Buraczyk, Anantaa Kotal

    Abstract: As machine learning systems move from theory to practice, they are increasingly tasked with decisions that affect healthcare access, financial opportunities, hiring, and public services. In these contexts, accuracy is only one piece of the puzzle - models must also be fair to different groups, protect individual privacy, and remain accountable to stakeholders. Achieving all three is difficult: dif… ▽ More

    Submitted 11 November, 2025; originally announced November 2025.

  11. arXiv:2510.11299  [pdf, ps, other

    cs.CR cs.DB

    How to Get Actual Privacy and Utility from Privacy Models: the k-Anonymity and Differential Privacy Families

    Authors: Josep Domingo-Ferrer, David Sánchez

    Abstract: Privacy models were introduced in privacy-preserving data publishing and statistical disclosure control with the promise to end the need for costly empirical assessment of disclosure risk. We examine how well this promise is kept by the main privacy models. We find they may fail to provide adequate protection guarantees because of problems in their definition or incur unacceptable trade-offs betwe… ▽ More

    Submitted 17 October, 2025; v1 submitted 13 October, 2025; originally announced October 2025.

    Comments: 13 pages

    MSC Class: 68 ACM Class: K.4.1

  12. arXiv:2510.11276  [pdf, ps, other

    physics.data-an cs.IT physics.app-ph

    Information-theoretic analysis of temporal dependence in discrete stochastic processes: Application to precipitation predictability

    Authors: Juan De Gregorio, David Sánchez, Raúl Toral

    Abstract: Understanding the temporal dependence of precipitation is key to improving weather predictability and developing efficient stochastic rainfall models. We introduce an information-theoretic approach to quantify memory effects in discrete stochastic processes and apply it to daily precipitation records across the contiguous United States. The method is based on the predictability gain, a quantity de… ▽ More

    Submitted 13 March, 2026; v1 submitted 13 October, 2025; originally announced October 2025.

    Journal ref: Chaos 36, 033124 (1-14) (2026)

  13. arXiv:2509.21560  [pdf, ps, other

    cs.SD

    Preserving Russek's "Summermood" Using Reality Check and a DeltaLab DL-4 Approximation

    Authors: Jeremy Hyrkas, Pablo Dodero Carrillo, Teresa Díaz de Cossio Sánchez

    Abstract: As a contribution towards ongoing efforts to maintain electroacoustic compositions for live performance, we present a collection of Pure Data patches to preserve and perform Antonio Russek's piece "Summermood" for bass flute and live electronics. The piece, originally written for the DeltaLab DL-4 delay rack unit, contains score markings specific to the DL-4. Here, we approximate the sound and uni… ▽ More

    Submitted 25 September, 2025; originally announced September 2025.

    Comments: 6 pages, 10 figures, Pure Data Max Conference 2025

    Journal ref: Proceedings and Programs of PdMaxCon25~ (2025) 55-60

  14. arXiv:2509.01379  [pdf, ps, other

    cs.CL

    WATCHED: A Web AI Agent Tool for Combating Hate Speech by Expanding Data

    Authors: Paloma Piot, Diego Sánchez, Javier Parapar

    Abstract: Online harms are a growing problem in digital spaces, putting user safety at risk and reducing trust in social media platforms. One of the most persistent forms of harm is hate speech. To address this, we need tools that combine the speed and scale of automated systems with the judgment and insight of human moderators. These tools should not only find harmful content but also explain their decisio… ▽ More

    Submitted 1 September, 2025; originally announced September 2025.

  15. arXiv:2507.04771  [pdf, ps, other

    cs.CR cs.LG

    Efficient Unlearning with Privacy Guarantees

    Authors: Josep Domingo-Ferrer, Najeeb Jebreel, David Sánchez

    Abstract: Privacy protection laws, such as the GDPR, grant individuals the right to request the forgetting of their personal data not only from databases but also from machine learning (ML) models trained on them. Machine unlearning has emerged as a practical means to facilitate model forgetting of data instances seen during training. Although some existing machine unlearning methods guarantee exact forgett… ▽ More

    Submitted 26 June, 2026; v1 submitted 7 July, 2025; originally announced July 2025.

    Comments: 34 pages, 9 tables, 2 figures

  16. arXiv:2504.18203  [pdf, other

    cs.CV cs.LG

    LiDAR-Guided Monocular 3D Object Detection for Long-Range Railway Monitoring

    Authors: Raul David Dominguez Sanchez, Xavier Diaz Ortiz, Xingcheng Zhou, Max Peter Ronecker, Michael Karner, Daniel Watzenig, Alois Knoll

    Abstract: Railway systems, particularly in Germany, require high levels of automation to address legacy infrastructure challenges and increase train traffic safely. A key component of automation is robust long-range perception, essential for early hazard detection, such as obstacles at level crossings or pedestrians on tracks. Unlike automotive systems with braking distances of ~70 meters, trains require pe… ▽ More

    Submitted 25 April, 2025; originally announced April 2025.

    Comments: Accepted for the Data-Driven Learning for Intelligent Vehicle Applications Workshop at the 36th IEEE Intelligent Vehicles Symposium (IV) 2025

  17. DP2Unlearning: An Efficient and Guaranteed Unlearning Framework for LLMs

    Authors: Tamim Al Mahmud, Najeeb Jebreel, Josep Domingo-Ferrer, David Sanchez

    Abstract: Large language models (LLMs) have recently revolutionized language processing tasks but have also brought ethical and legal issues. LLMs have a tendency to memorize potentially private or copyrighted information present in the training data, which might then be delivered to end users at inference time. When this happens, a naive solution is to retrain the model from scratch after excluding the und… ▽ More

    Submitted 18 July, 2025; v1 submitted 18 April, 2025; originally announced April 2025.

    Comments: This is the updated version of the preprint, revised following acceptance for publication in Elsevier Neural Networks Journal. The paper is now published (18 July 2025) with DOI: https://doi.org/10.1016/j.neunet.2025.107879

    Journal ref: Neural Networks, 2025, Article 107879

  18. arXiv:2503.08188  [pdf, other

    cs.CL cs.AI

    RigoChat 2: an adapted language model to Spanish using a bounded dataset and reduced hardware

    Authors: Gonzalo Santamaría Gómez, Guillem García Subies, Pablo Gutiérrez Ruiz, Mario González Valero, Natàlia Fuertes, Helena Montoro Zamorano, Carmen Muñoz Sanz, Leire Rosado Plaza, Nuria Aldama García, David Betancur Sánchez, Kateryna Sushkova, Marta Guerrero Nieto, Álvaro Barbero Jiménez

    Abstract: Large Language Models (LLMs) have become a key element of modern artificial intelligence, demonstrating the ability to address a wide range of language processing tasks at unprecedented levels of accuracy without the need of collecting problem-specific data. However, these versatile models face a significant challenge: both their training and inference processes require substantial computational r… ▽ More

    Submitted 11 March, 2025; originally announced March 2025.

  19. arXiv:2501.16011  [pdf, other

    cs.CL

    MEL: Legal Spanish Language Model

    Authors: David Betancur Sánchez, Nuria Aldama García, Álvaro Barbero Jiménez, Marta Guerrero Nieto, Patricia Marsà Morales, Nicolás Serrano Salas, Carlos García Hernán, Pablo Haya Coll, Elena Montiel Ponsoda, Pablo Calleja Ibáñez

    Abstract: Legal texts, characterized by complex and specialized terminology, present a significant challenge for Language Models. Adding an underrepresented language, such as Spanish, to the mix makes it even more challenging. While pre-trained models like XLM-RoBERTa have shown capabilities in handling multilingual corpora, their performance on domain specific documents remains underexplored. This paper pr… ▽ More

    Submitted 27 January, 2025; originally announced January 2025.

    Comments: 8 pages, 6 figures, 3 tables

  20. arXiv:2501.15990  [pdf, other

    cs.CL

    3CEL: A corpus of legal Spanish contract clauses

    Authors: Nuria Aldama García, Patricia Marsà Morales, David Betancur Sánchez, Álvaro Barbero Jiménez, Marta Guerrero Nieto, Pablo Haya Coll, Patricia Martín Chozas, Elena Montiel Ponsoda

    Abstract: Legal corpora for Natural Language Processing (NLP) are valuable and scarce resources in languages like Spanish due to two main reasons: data accessibility and legal expert knowledge availability. INESData 2024 is a European Union funded project lead by the Universidad Politécnica de Madrid (UPM) and developed by Instituto de Ingeniería del Conocimiento (IIC) to create a series of state-of-the-art… ▽ More

    Submitted 27 January, 2025; originally announced January 2025.

    Comments: 12 pages, 13 figures, 6 tables

  21. arXiv:2501.00563  [pdf, other

    math.AG cs.SC math.KT

    Motives meet SymPy: studying $λ$-ring expressions in Python

    Authors: Daniel Sanchez, David Alfaya, Jaime Pizarroso

    Abstract: We present a new Python package called "motives", a symbolic manipulation package based on SymPy capable of handling and simplifying motivic expressions in the Grothendieck ring of Chow motives and other types of $λ$-rings. The package is able to manipulate and compare arbitrary expressions in $λ$-rings and, in particular, it contains explicit tools for manipulating motives of several types of com… ▽ More

    Submitted 31 December, 2024; originally announced January 2025.

    Comments: 19 pages, 2 figures. The code of the library described in the paper is publicly hosted at https://github.com/CIAMOD/motives

    MSC Class: 13D15 (Primary) 68W30; 19E08; 14C35; 14D20; 14H60 (Secondary) ACM Class: I.1.1

  22. arXiv:2412.12928  [pdf, ps, other

    cs.CL

    Truthful Text Sanitization Guided by Inference Attacks

    Authors: Ildikó Pilán, Benet Manzanares-Salor, David Sánchez, Pierre Lison

    Abstract: Text sanitization aims to rewrite parts of a document to prevent disclosure of personal information. The central challenge of text sanitization is to strike a balance between privacy protection (avoiding the leakage of personal information) and utility preservation (retaining as much as possible of the document's original content). To this end, we introduce a novel text sanitization method based o… ▽ More

    Submitted 31 August, 2025; v1 submitted 17 December, 2024; originally announced December 2024.

  23. arXiv:2412.05385  [pdf

    cs.NI eess.SY

    Enhanced 5G/B5G Network Planning/Optimization deploying RIS in Urban/Outdoor Scenarios

    Authors: Valdemar Farré, Juan C. Estrada-Jiménez, José D. Vega Sánchez, Juan A. Vasquez-Peralvo, Symeon Chatzinotas

    Abstract: In recent years, the fifth-generation (5G) mobile network has been developed worldwide to remarkably improve network performance and spectral efficiency. Very recently, reconfigurable intelligent surfaces (RISs) technology has emerged as an innovative solution for controlling the propagation medium of the forthcoming sixth-generation (6G) networks. Specifically, RIS takes advantage of the reflecte… ▽ More

    Submitted 6 December, 2024; originally announced December 2024.

    Comments: 6 pages, 12 figures, 2024 IEEE 29th International Workshop on Computer Aided Modeling and Design of Communication Links and Networks (CAMAD), presented in Session 11: Federation of 6G Infrastructures and Experimentation Facilities in a Network of Networks, October 22th - 2024

  24. arXiv:2411.10227  [pdf, ps, other

    cs.CL cs.IR physics.soc-ph

    Entropy and type-token ratio in gigaword corpora

    Authors: Pablo Rosillo-Rodes, Maxi San Miguel, David Sanchez

    Abstract: There are different ways of measuring diversity in complex systems. In particular, in language, lexical diversity is characterized in terms of the type-token ratio and the word entropy. We here investigate both diversity metrics in six massive linguistic datasets in English, Spanish, and Turkish, consisting of books, news articles, and tweets. These gigaword corpora correspond to languages with di… ▽ More

    Submitted 24 June, 2025; v1 submitted 15 November, 2024; originally announced November 2024.

    Comments: 15 pages, 10 figures, 8 tables

    Journal ref: Phys. Rev. Research 7, 033054 (2025)

  25. arXiv:2405.05723  [pdf, ps, other

    cs.CL cs.AI cs.IR

    Computational lexical analysis of Flamenco genres

    Authors: Pablo Rosillo-Rodes, Maxi San Miguel, David Sanchez

    Abstract: Flamenco, recognized by UNESCO as part of the Intangible Cultural Heritage of Humanity, is a profound expression of cultural identity rooted in Andalusia, Spain. However, there is a lack of quantitative studies that help identify characteristic patterns in this long-lived music tradition. In this work, we present a computational analysis of Flamenco lyrics, employing natural language processing an… ▽ More

    Submitted 13 March, 2026; v1 submitted 9 May, 2024; originally announced May 2024.

    Comments: 25 pages, 20 figures

    Journal ref: ACM J. Comput. Cult. Herit. 18, 59 (2025)

  26. Digital Forgetting in Large Language Models: A Survey of Unlearning Methods

    Authors: Alberto Blanco-Justicia, Najeeb Jebreel, Benet Manzanares, David Sánchez, Josep Domingo-Ferrer, Guillem Collell, Kuan Eeik Tan

    Abstract: The objective of digital forgetting is, given a model with undesirable knowledge or behavior, obtain a new model where the detected issues are no longer present. The motivations for forgetting include privacy protection, copyright protection, elimination of biases and discrimination, and prevention of harmful content generation. Effective digital forgetting has to be effective (meaning how well th… ▽ More

    Submitted 2 April, 2024; originally announced April 2024.

    Comments: 70 pages

    MSC Class: 68 ACM Class: K.4.1; I.2.6; I.2.7

    Journal ref: Artificial Intelligence Review, vol. 58, art. no. 90, 2025

  27. arXiv:2403.18430  [pdf, other

    cs.CL physics.data-an physics.soc-ph stat.AP

    Exploring language relations through syntactic distances and geographic proximity

    Authors: Juan De Gregorio, Raúl Toral, David Sánchez

    Abstract: Languages are grouped into families that share common linguistic traits. While this approach has been successful in understanding genetic relations between diverse languages, more analyses are needed to accurately quantify their relatedness, especially in less studied linguistic levels such as syntax. Here, we explore linguistic distances using series of parts of speech (POS) extracted from the Un… ▽ More

    Submitted 3 October, 2024; v1 submitted 27 March, 2024; originally announced March 2024.

    Comments: 39 pages

    Journal ref: EPJ Data Science 13, 61 (2024)

  28. arXiv:2402.09275  [pdf

    cs.SI cs.CY physics.soc-ph

    The socialisation of the adolescent who carries out team sports: a transversal study of centrality with a social network analysis

    Authors: Pilar Marqués-Sánchez, José Alberto Benítez-Andrades, María Dolores Calvo Sánchez, Natalia Arias

    Abstract: Objectives: This study analyzed adolescent physical activity, its link to overweight, and the social network structure in group sports participants, focusing on centrality measures. Setting: Conducted in 11 classrooms across 5 schools in Ponferrada, Spain. Participants: Included 235 adolescents (49.4% female), categorized as normal weight or overweight. Methods: The Physical Activity Questio… ▽ More

    Submitted 14 February, 2024; originally announced February 2024.

    Journal ref: BMJ Open, 2021, Volume 11, ID e042773

  29. arXiv:2312.13712  [pdf, other

    cs.CR

    Conciliating Privacy and Utility in Data Releases via Individual Differential Privacy and Microaggregation

    Authors: Jordi Soria-Comas, David Sánchez, Josep Domingo-Ferrer, Sergio Martínez, Luis Del Vasto-Terrientes

    Abstract: $ε$-Differential privacy (DP) is a well-known privacy model that offers strong privacy guarantees. However, when applied to data releases, DP significantly deteriorates the analytical utility of the protected outcomes. To keep data utility at reasonable levels, practical applications of DP to data releases have used weak privacy parameters (large $ε… ▽ More

    Submitted 21 December, 2023; originally announced December 2023.

    Comments: 17 pages, 6 figures

  30. arXiv:2311.11882  [pdf, other

    cs.CV cs.LG

    Multi-Task Faces (MTF) Data Set: A Legally and Ethically Compliant Collection of Face Images for Various Classification Tasks

    Authors: Rami Haffar, David Sánchez, Josep Domingo-Ferrer

    Abstract: Human facial data offers valuable potential for tackling classification problems, including face recognition, age estimation, gender identification, emotion analysis, and race classification. However, recent privacy regulations, particularly the EU General Data Protection Regulation, have restricted the collection and usage of human images in research. As a result, several previously published fac… ▽ More

    Submitted 8 April, 2025; v1 submitted 20 November, 2023; originally announced November 2023.

    Comments: 21 pages, 2 figures, 9 Tables,

  31. arXiv:2311.03171  [pdf, other

    cs.CR cs.LG

    An Examination of the Alleged Privacy Threats of Confidence-Ranked Reconstruction of Census Microdata

    Authors: David Sánchez, Najeeb Jebreel, Krishnamurty Muralidhar, Josep Domingo-Ferrer, Alberto Blanco-Justicia

    Abstract: The threat of reconstruction attacks has led the U.S. Census Bureau (USCB) to replace in the Decennial Census 2020 the traditional statistical disclosure limitation based on rank swapping with one based on differential privacy (DP), leading to substantial accuracy loss of released statistics. Yet, it has been argued that, if many different reconstructions are compatible with the released statistic… ▽ More

    Submitted 17 September, 2024; v1 submitted 6 November, 2023; originally announced November 2023.

    Comments: In Lecture Notes in Artificial Intelligence, vol. 14915, pp. 213-224. Vol. Privacy in Statistical Databases (PSD 2024), Antibes Juan-les-Pins, France, Sep. 25-27, 2024. 20 pages, 5 figures, 4 tables

  32. arXiv:2307.10016  [pdf, ps, other

    physics.soc-ph cs.CL cs.SI

    When Dialects Collide: How Socioeconomic Mixing Affects Language Use

    Authors: Thomas Louf, José J. Ramasco, David Sánchez, Márton Karsai

    Abstract: The socioeconomic background of people and how they use standard forms of language are not independent, as demonstrated in various sociolinguistic studies. However, the extent to which these correlations may be influenced by the mixing of people from different socioeconomic classes remains relatively unexplored from a quantitative perspective. In this work we leverage geotagged tweets and transfer… ▽ More

    Submitted 10 July, 2025; v1 submitted 19 July, 2023; originally announced July 2023.

    Journal ref: EPJ Data Sci. 14, 47 (2025)

  33. arXiv:2212.02448  [pdf, ps, other

    cs.IT

    The Multi-cluster Fluctuating Two-Ray Fading Model

    Authors: José David Vega Sánchez, F. Javier López-Martínez, José F. Paris, Juan M. Romero-Jerez

    Abstract: We introduce a new class of fading channels, built as the superposition of two fluctuating specular components with random phases, plus a clustering of scattered waves: the Multi-cluster Fluctuating Two-Ray (MFTR) fading channel. The MFTR model emerges as a natural generalization of both the fluctuating two-ray (FTR) and the $κ$-$μ$ shadowed fading models through a more general yet equally mathema… ▽ More

    Submitted 15 September, 2023; v1 submitted 5 December, 2022; originally announced December 2022.

    Comments: This work was submitted to the IEEE for publication on May 31, 2022

  34. arXiv:2210.06139  [pdf, other

    cs.CR cs.CE econ.GN

    Zero-Knowledge Optimal Monetary Policy under Stochastic Dominance

    Authors: David Cerezo Sánchez

    Abstract: Optimal simple rules for the monetary policy of the first stochastically dominant crypto-currency are derived in a Dynamic Stochastic General Equilibrium (DSGE) model, in order to provide optimal responses to changes in inflation, output, and other sources of uncertainty. The optimal monetary policy stochastically dominates all the previous crypto-currencies, thus the efficient portfolio is to go… ▽ More

    Submitted 12 October, 2022; originally announced October 2022.

    Comments: Implementation available at: https://github.com/Calctopia-OpenSource/cothority/tree/zkmonpolicy

  35. arXiv:2209.06375  [pdf, other

    cs.CV astro-ph.IM

    Self-Supervised Clustering on Image-Subtracted Data with Deep-Embedded Self-Organizing Map

    Authors: Y. -L. Mong, K. Ackley, T. L. Killestein, D. K. Galloway, M. Dyer, R. Cutter, M. J. I. Brown, J. Lyman, K. Ulaczyk, D. Steeghs, V. Dhillon, P. O'Brien, G. Ramsay, K. Noysena, R. Kotak, R. Breton, L. Nuttall, E. Palle, D. Pollacco, E. Thrane, S. Awiphan, U. Burhanudin, P. Chote, A. Chrimes, E. Daw , et al. (23 additional authors not shown)

    Abstract: Developing an effective automatic classifier to separate genuine sources from artifacts is essential for transient follow-ups in wide-field optical surveys. The identification of transient detections from the subtraction artifacts after the image differencing process is a key step in such classifiers, known as real-bogus classification problem. We apply a self-supervised machine learning model, th… ▽ More

    Submitted 13 September, 2022; originally announced September 2022.

  36. arXiv:2209.03956  [pdf, other

    q-bio.NC cs.AI

    Technology and Consciousness

    Authors: John Rushby, Daniel Sanchez

    Abstract: We report on a series of eight workshops held in the summer of 2017 on the topic "technology and consciousness." The workshops covered many subjects but the overall goal was to assess the possibility of machine consciousness, and its potential implications. In the body of the report, we summarize most of the basic themes that were discussed: the structure and function of the brain, theories of con… ▽ More

    Submitted 17 July, 2022; originally announced September 2022.

    Comments: SRI CSL Workshop Report 2017-1

  37. arXiv:2208.11175  [pdf, other

    cs.CL cond-mat.stat-mech physics.soc-ph

    Ordinal analysis of lexical patterns

    Authors: David Sanchez, Luciano Zunino, Juan De Gregorio, Raul Toral, Claudio Mirasso

    Abstract: Words are fundamental linguistic units that connect thoughts and things through meaning. However, words do not appear independently in a text sequence. The existence of syntactic rules induces correlations among neighboring words. Using an ordinal pattern approach, we present an analysis of lexical statistical connections for 11 major languages. We find that the diverse manners that languages util… ▽ More

    Submitted 14 March, 2023; v1 submitted 23 August, 2022; originally announced August 2022.

    Comments: 9 pages, 12 figures, 2 tables; v2: the section on universality has been removed because previous results were affected by spurious correlations. Published version

    Journal ref: Chaos 33, 033121 (2023)

  38. arXiv:2208.07649  [pdf, other

    cs.CL cs.CY cs.SI physics.soc-ph

    American cultural regions mapped through the lexical analysis of social media

    Authors: Thomas Louf, Bruno Gonçalves, Jose J. Ramasco, David Sanchez, Jack Grieve

    Abstract: Cultural areas represent a useful concept that cross-fertilizes diverse fields in social sciences. Knowledge of how humans organize and relate their ideas and behavior within a society helps to understand their actions and attitudes towards different issues. However, the selection of common traits that shape a cultural area is somewhat arbitrary. What is needed is a method that can leverage the ma… ▽ More

    Submitted 18 April, 2023; v1 submitted 16 August, 2022; originally announced August 2022.

    Comments: 13 pages, 5 figures; contains Supplementary Information

    Journal ref: Humanit Soc Sci Commun 10, 133 (2023)

  39. Enhanced Security and Privacy via Fragmented Federated Learning

    Authors: Najeeb Moharram Jebreel, Josep Domingo-Ferrer, Alberto Blanco-Justicia, David Sanchez

    Abstract: In federated learning (FL), a set of participants share updates computed on their local data with an aggregator server that combines updates into a global model. However, reconciling accuracy with privacy and security is a challenge to FL. On the one hand, good updates sent by honest participants may reveal their private local information, whereas poisoned updates sent by malicious participants ma… ▽ More

    Submitted 19 November, 2022; v1 submitted 13 July, 2022; originally announced July 2022.

    Comments: IEEE Transactions on Neural Networks and Learning Systems (To Appear)

  40. arXiv:2207.01982  [pdf, other

    cs.CR cs.LG

    Defending against the Label-flipping Attack in Federated Learning

    Authors: Najeeb Moharram Jebreel, Josep Domingo-Ferrer, David Sánchez, Alberto Blanco-Justicia

    Abstract: Federated learning (FL) provides autonomy and privacy by design to participating peers, who cooperatively build a machine learning (ML) model while keeping their private data in their devices. However, that same autonomy opens the door for malicious peers to poison the model by conducting either untargeted or targeted poisoning attacks. The label-flipping (LF) attack is a targeted poisoning attack… ▽ More

    Submitted 5 July, 2022; originally announced July 2022.

  41. arXiv:2206.04621  [pdf, ps, other

    cs.CR cs.LG

    A Critical Review on the Use (and Misuse) of Differential Privacy in Machine Learning

    Authors: Alberto Blanco-Justicia, David Sanchez, Josep Domingo-Ferrer, Krishnamurty Muralidhar

    Abstract: We review the use of differential privacy (DP) for privacy protection in machine learning (ML). We show that, driven by the aim of preserving the accuracy of the learned models, DP-based ML implementations are so loose that they do not offer the ex ante privacy guarantees of DP. Instead, what they deliver is basically noise addition similar to the traditional (and often criticized) statistical dis… ▽ More

    Submitted 5 July, 2022; v1 submitted 9 June, 2022; originally announced June 2022.

    Comments: ACM Computing Surveys (to appear)

    ACM Class: I.2.6

    Journal ref: ACM Computing Surveys, vol. 55, no. 8, pp. 1-26, 2023

  42. arXiv:2205.10233  [pdf, other

    cs.CL

    RigoBERTa: A State-of-the-Art Language Model For Spanish

    Authors: Alejandro Vaca Serrano, Guillem Garcia Subies, Helena Montoro Zamorano, Nuria Aldama Garcia, Doaa Samy, David Betancur Sanchez, Antonio Moreno Sandoval, Marta Guerrero Nieto, Alvaro Barbero Jimenez

    Abstract: This paper presents RigoBERTa, a State-of-the-Art Language Model for Spanish. RigoBERTa is trained over a well-curated corpus formed up from different subcorpora with key features. It follows the DeBERTa architecture, which has several advantages over other architectures of similar size as BERT or RoBERTa. RigoBERTa performance is assessed over 13 NLU tasks in comparison with other available Spani… ▽ More

    Submitted 3 June, 2022; v1 submitted 27 April, 2022; originally announced May 2022.

  43. arXiv:2203.08370  [pdf, ps, other

    cs.IT

    Physical Layer Security of RIS-Assisted Communications under Electromagnetic Interference

    Authors: José David Vega Sánchez, Georges Kaddoum, F. Javier López-Martínez

    Abstract: This work investigates the impact of the ever-present electromagnetic interference (EMI) on the achievable secrecy performance of reconfigurable intelligent surface (RIS)-aided communication systems. We characterize the end-to-end RIS channel by considering key practical aspects such as spatial correlation, transmit beamforming vector, phase-shift noise, the coexistence of direct and indirect chan… ▽ More

    Submitted 15 March, 2022; originally announced March 2022.

  44. arXiv:2202.00443  [pdf, other

    cs.CL cs.AI

    The Text Anonymization Benchmark (TAB): A Dedicated Corpus and Evaluation Framework for Text Anonymization

    Authors: Ildikó Pilán, Pierre Lison, Lilja Øvrelid, Anthi Papadopoulou, David Sánchez, Montserrat Batet

    Abstract: We present a novel benchmark and associated evaluation metrics for assessing the performance of text anonymization methods. Text anonymization, defined as the task of editing a text document to prevent the disclosure of personal information, currently suffers from a shortage of privacy-oriented annotated text resources, making it difficult to properly evaluate the level of privacy protection offer… ▽ More

    Submitted 1 July, 2022; v1 submitted 25 January, 2022; originally announced February 2022.

  45. arXiv:2109.07267  [pdf, other

    cs.CR cs.GT econ.GN

    JUBILEE: Secure Debt Relief and Forgiveness

    Authors: David Cerezo Sánchez

    Abstract: JUBILEE is a securely computed mechanism for debt relief and forgiveness in a frictionless manner without involving trusted third parties, leading to more harmonious debt settlements by incentivising the parties to truthfully reveal their private information. JUBILEE improves over all previous methods: - individually rational, incentive-compatible, truthful/strategy-proof, ex-post efficient, opt… ▽ More

    Submitted 15 September, 2021; originally announced September 2021.

  46. arXiv:2109.05371  [pdf, other

    cs.CR cs.AR

    F1: A Fast and Programmable Accelerator for Fully Homomorphic Encryption (Extended Version)

    Authors: Axel Feldmann, Nikola Samardzic, Aleksandar Krastev, Srini Devadas, Ron Dreslinski, Karim Eldefrawy, Nicholas Genise, Chris Peikert, Daniel Sanchez

    Abstract: Fully Homomorphic Encryption (FHE) allows computing on encrypted data, enabling secure offloading of computation to untrusted serves. Though it provides ideal security, FHE is expensive when executed in software, 4 to 5 orders of magnitude slower than computing on unencrypted data. These overheads are a major barrier to FHE's widespread adoption. We present F1, the first FHE accelerator that is pr… ▽ More

    Submitted 25 September, 2021; v1 submitted 11 September, 2021; originally announced September 2021.

  47. arXiv:2108.01913  [pdf, other

    cs.CR cs.DC cs.GT cs.LG

    Secure and Privacy-Preserving Federated Learning via Co-Utility

    Authors: Josep Domingo-Ferrer, Alberto Blanco-Justicia, Jesús Manjón, David Sánchez

    Abstract: The decentralized nature of federated learning, that often leverages the power of edge devices, makes it vulnerable to attacks against privacy and security. The privacy risk for a peer is that the model update she computes on her private data may, when sent to the model manager, leak information on those private data. Even more obvious are security attacks, whereby one or several malicious peers r… ▽ More

    Submitted 4 August, 2021; originally announced August 2021.

    Comments: IEEE Internet of Things Journal, to appear

    MSC Class: 68P27; 68Txx; 91 ACM Class: I.2.11; K.6.5

  48. arXiv:2105.10464  [pdf, other

    cs.CR cs.DC econ.GN

    Pravuil: Global Consensus for a United World

    Authors: David Cerezo Sánchez

    Abstract: Pravuil is a robust, secure, and scalable consensus protocol for a permissionless blockchain suitable for deployment in an adversarial environment such as the Internet. Pravuil circumvents previous shortcomings of other blockchains: - Bitcoin's limited adoption problem: as transaction demand grows, payment confirmation times grow much lower than other PoW blockchains - higher transaction secur… ▽ More

    Submitted 21 May, 2021; originally announced May 2021.

    Journal ref: FinTech 2022, 1(4), 325-344

  49. arXiv:2105.02570  [pdf, other

    physics.soc-ph cs.CL cs.SI

    Capturing the diversity of multilingual societies

    Authors: Thomas Louf, David Sanchez, Jose J. Ramasco

    Abstract: Cultural diversity encoded within languages of the world is at risk, as many languages have become endangered in the last decades in a context of growing globalization. To preserve this diversity, it is first necessary to understand what drives language extinction, and which mechanisms might enable coexistence. Here, we study language shift mechanisms using theoretical and data-driven perspectives… ▽ More

    Submitted 7 October, 2022; v1 submitted 6 May, 2021; originally announced May 2021.

    Comments: Main text: 12 pages, 6 figures, 51 references. Supplementary Information: 27 pages, 16 figures, 2 tables

    Journal ref: Phys. Rev. Research 3, 043146 (2021)

  50. arXiv:2103.13525  [pdf, ps, other

    cs.IT

    Expectation-Maximization Learning for Wireless Channel Modeling of Reconfigurable Intelligent Surfaces

    Authors: José David Vega Sánchez, Luis Urquiza-Aguiar, Martha Cecilia Paredes Paredes, F. Javier López-Martínez

    Abstract: Channel modeling is a critical issue when designing or evaluating the performance of reconfigurable intelligent surface (RIS)-assisted communications. Inspired by the promising potential of learning-based methods for characterizing the radio environment, we present a general approach to model the RIS end-to-end equivalent channel using the unsupervised expectation-maximization (EM) learning algori… ▽ More

    Submitted 10 August, 2021; v1 submitted 24 March, 2021; originally announced March 2021.