Sorry, you need to enable JavaScript to visit this website.
Partager

Publications

 

Les publications de nos enseignants-chercheurs sont sur la plateforme HAL :

 

Les publications des thèses des docteurs du LTCI sont sur la plateforme HAL :

 

Retrouver les publications figurant dans l'archive ouverte HAL par année :

2026

  • Actualizing Ethical Principles for Curating Large-Scale Training Datasets in the Era of Massive AI Models
    • Cazacu Silvia
    • Qian Alice
    • Zhao Dora
    • Pine Kathleen
    • Walker Shawn
    • Shen Hong
    • Dabbish Laura
    • Panagiotidou Georgia
    • Klaus Scheuerman Morgan
    , 2026, 21. While AI technologies are often framed as ubiquitous and inevitable, their expansion relies on the large-scale extraction of data from diverse global communities. However, the datasets powering foundation models are often treated as found artifacts rather than products of specific power dynamics and human labor. Current practices frequently disregard the structural inequities embedded in data, even as these systems profoundly impact systemically marginalized communities. While frameworks for ethical curation exist for smaller datasets, the massive scale of foundation models has introduced a logic of extraction that prioritizes volume over accountability. This workshop invites researchers, practitioners, and activists to move beyond standard technical hurdles and instead problematize the foundational assumptions of large-scale data work. This workshop builds on a series of ongoing conversations across scholarly communities and continues work initiated at CSCW 2025. Drawing from the CRAFT tradition of transdisciplinary exchange and community action, we will facilitate a collective refactoring of the three core pillars of data curation: composition, focused on interrogating whose lives are extracted and how representation is shaped by hegemonic interests; process, which centers the invisible labor and situated contexts involved in curating and cleaning massive data stores, and release, focused on rethinking the governance, accountability, and potential for refusal in how these models are shared with the world. Our goal is to cultivate a community-led conceptual framework that reimagines more just sociotechnical futures—transforming data curation from a top-down technical requirement into an act of collective responsibility and repair. (10.1108/jices-08-2022-0069)
    DOI : 10.1108/jices-08-2022-0069
  • RAVE: Rate-Adaptive Visual Encoding for 3D Gaussian Splatting
    • Tran Hoang-Nhat
    • Di Sario Francesco
    • Spadaro Gabriele
    • Valenzise Giuseppe
    • Tartaglione Enzo
    , 2026, pp.11727-11731. <div><p>Recent advances in neural scene representations have transformed immersive multimedia, with 3D Gaussian Splatting (3DGS) enabling real-time photorealistic rendering. Despite its efficiency, 3DGS suffers from large memory requirements and costly training procedures, motivating efforts toward compression. Existing approaches, however, operate at fixed rates, limiting adaptability to varying bandwidth and device constraints. In this work, we propose a flexible compression scheme for 3DGS that supports interpolation at any rate between predefined bounds. Our method is computationally lightweight, requires no retraining for any rate, and preserves rendering quality across a broad range of operating points. Experiments demonstrate that the approach achieves efficient, high-quality compression while offering dynamic rate control, making it suitable for practical deployment in immersive applications. The code is available at https://github.com/inspiros/RAVE.</p></div> (10.1109/ICASSP55912.2026.11463333)
    DOI : 10.1109/ICASSP55912.2026.11463333
  • S-PRESSO: Ultra Low Bitrate Sound Effect Compression With Diffusion Autoencoders And Offline Quantization
    • Lahrichi Zineb
    • Hadjeres Gaëtan
    • Richard Gaël
    • Peeters Geoffroy
    , 2026. <div><p>Neural audio compression models have recently achieved extreme compression rates, enabling efficient latent generative modeling. Conversely, latent generative models have been applied to compression, pushing the limits of continuous and discrete approaches. However, existing methods remain constrained to low-resolution audio and degrade substantially at very low bitrates, where audible artifacts are prominent. In this paper, we present S-PRESSO, a 48kHz sound effect compression model that produces both continuous and discrete embeddings at ultra-low bitrates, down to 0.096 kbps, via offline quantization. Our model relies on a pretrained latent diffusion model to decode compressed audio embeddings learned by a latent encoder. Leveraging the generative priors of the diffusion decoder, we achieve extremely low frame rates, down to 1Hz (750x compression rate), producing convincing and realistic reconstructions at the cost of exact fidelity. Despite operating at high compression rates, we demonstrate that S-PRESSO outperforms both continuous and discrete baselines in audio quality, acoustic similarity and reconstruction metrics.</p></div>
  • HFMCA: Orthonormal Feature Learning for EEG-Based Brain Decoding
    • Wang Yinghao
    • Xu Lintao
    • Yu Shujian
    • Tartaglione Enzo
    • Nguyen Van-Tam
    , 2026, pp.6826-6830. Electroencephalography (EEG) analysis is critical for brain-computer interfaces and neuroscience, but the intrinsic noise and high dimensionality of EEG signals hinder effective feature learning. We propose a self-supervised framework based on the Hierarchical Functional Maximal Correlation Algorithm (HFMCA), which learns orthonormal EEG representations by enforcing feature decorrelation and reducing redundancy. This design enables robust capture of essential brain dynamics for various EEG recognition tasks. We validate HFMCA on two benchmark datasets, SEED and BCIC-2A, where pretraining with HFMCA consistently outperforms competitive self-supervised baselines, achieving notable gains in classification accuracy. Across diverse EEG tasks, our method demonstrates superior cross-subject generalization under leave-one-subject-out validation, advancing state-of-the-art by 2.71% on SEED emotion recognition and 2.57% on BCIC-2A motor imagery classification. Our code and supplementary material are available at: https://github.com/W-Yinghao/HFMCA_EEG. (10.1109/ICASSP55912.2026.11464747)
    DOI : 10.1109/ICASSP55912.2026.11464747
  • SIRUP: A DIFFUSION-BASED VIRTUAL UPMIXER OF STEERING VECTORS FOR HIGHLY-DIRECTIVE SPATIALIZATION WITH FIRST-ORDER AMBISONICS
    • Picard Emilio
    • Carlo Diego Di
    • Nugraha Aditya Arie
    • Fontaine Mathieu
    • Yoshii Kazuyoshi
    , 2026, pp.14707-14711. <div><p>This paper presents virtual upmixing of steering vectors captured by a fewer-channel spherical microphone array. This challenge has conventionally been addressed by recovering the directions and signals of sound sources from first-order ambisonics (FOA) data, and then rendering the higher-order ambisonics (HOA) data using a physics-based acoustic simulator. This approach, however, struggles to handle the mutual dependency between the spatial directivity of source estimation and the spatial resolution of FOA ambisonics data. Our method, named SIRUP, employs a latent diffusion model architecture. Specifically, a variational autoencoder (VAE) is used to learn a compact encoding of the HOA data in a latent space and a diffusion model is then trained to generate the HOA embeddings, conditioned by the FOA data. Experimental results showed that SIRUP achieved a significant improvement compared to FOA systems for steering vector upmixing, source localization, and speech denoising.</p></div> (10.1109/ICASSP55912.2026.11464234)
    DOI : 10.1109/ICASSP55912.2026.11464234
  • Residual Tokens Enhance Masked Autoencoders For Speech Modeling
    • Sadok Samir
    • Lathuilière Stéphane
    • Alameda-Pineda Xavier
    , 2026, pp.14447-14451. Recent speech modeling relies on explicit attributes such as pitch, content, and speaker identity, but these alone cannot capture the full richness of natural speech. We introduce RT-MAE, a novel masked autoencoder framework that augments the supervised attributes-based modeling with unsupervised residual trainable tokens, designed to encode the information not explained by explicit labeled factors (e.g., timbre variations, noise, emotion etc). Experiments show that RT-MAE improves reconstruction quality, preserving content and speaker similarity while enhancing expressivity. We further demonstrate its applicability to speech enhancement, removing noise at inference while maintaining controllability and naturalness. (10.1109/ICASSP55912.2026.11464180)
    DOI : 10.1109/ICASSP55912.2026.11464180
  • GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech Synthesis
    • Baoueb Teysir
    • Bie Xiaoyu
    • Fontaine Mathieu
    • Richard Gaël
    , 2026, pp.17732-17736. <div><p>Recent advances in diffusion models have positioned them as powerful generative frameworks for speech synthesis, demonstrating substantial improvements in audio quality and stability. Nevertheless, their effectiveness in vocoders conditioned on mel spectrograms remains constrained, particularly when the conditioning diverges from the training distribution. The recently proposed GLA-Grad model introduced a phase-aware extension to the WaveGrad vocoder that integrated the Griffin-Lim algorithm (GLA) into the reverse process to reduce inconsistencies between generated signals and conditioning mel spectrogram. In this paper, we further improve GLA-Grad through an innovative choice in how to apply the correction. Particularly, we compute the correction term only once, with a single application of GLA, to accelerate the generation process. Experimental results demonstrate that our method consistently outperforms the baseline models, particularly in out-of-domain scenarios.</p></div>
  • PHYSICS-INFORMED LEARNING OF NEURAL SCATTERING FIELDS TOWARDS MEASUREMENT-FREE MESH-TO-HRTF ESTIMATION
    • Martinez Tancrède
    • Carlo Diego Di
    • Nugraha Aditya Arie
    • Fontaine Mathieu
    • Yoshii Kazuyoshi
    , 2026, pp.22577-22581. <div><p>This paper describes neural simulation of the scattered pressure field from a plane wave around a scattering object in both continuous 2D and 3D domains. This task has typically been treated as a regression problem that aims to train a physicsinformed neural network (PINN) using pressure measurements at discrete positions. This approach, however, needs to train the whole network for each incident wave direction. To address this, we propose a measurement-free simulator based on a PINN purely driven by the Helmholtz equation with the Robin boundary condition and the Sommerfeld radiation condition with the aid of the perfectly matched layer (PML) framework. More specifically, we design a physics-informed scattering hypernetwork (PHISK) that can generalize to incident waves from any direction via low-rank adaptation (LoRA) of a PINN trained for a specific configuration. The experiment shows that the proposed method accurately simulated sound scattering around various objects, adapting to unseen incident wave directions with minimal performance loss, and realized reasonable simulation of head-related transfer functions (HRTFs) from complex mesh data of a human head.</p></div> (10.1109/ICASSP55912.2026.11462698)
    DOI : 10.1109/ICASSP55912.2026.11462698
  • Generalization Bounds for Spectral GNNs via Fourier Domain Analysis
    • Martirosyan Vahan A
    • Malitesta Daniele
    • Talbot Hugues
    • Giraldo Jhony H
    • Malliaros Fragkiskos D
    , 2026, 300. Spectral graph neural networks learn graph filters, but their behavior with increasing depth and polynomial order is not well understood. We analyze these models in the graph Fourier domain, where each layer becomes an element-wise frequency update, separating the fixed spectrum from trainable parameters and making depth and order explicit. In this setting, we show that Gaussian complexity is invariant under the Graph Fourier Transform, which allows us to derive data-dependent, depth, and order-aware generalization bounds together with stability estimates. In the linear case, our bounds are tighter, and on real graphs, the data-dependent term correlates with the generalization gap across polynomial bases, highlighting practical choices that avoid frequency amplification across layers.
  • Numerical Model of a Multiple-Input-Multiple-Output Distributed Acoustic Sensor with Joint Phase and Birefringence Estimation
    • Prato Diane
    • Sheramin Mehran Mokhtari
    • Gabet Renaud
    • Awwad Élie
    , 2026. In this work, we introduce and experimentally validate a numerical model for a Multiple-Input-Multiple-Output Distributed Acoustic Sensing (MIMO-DAS) system that accounts for dynamic perturbations of fiber birefringence and of the common optical phase of the backscattered signal (or polarization-averaged phase, shared by both polarization tributaries). The MIMO-DAS system probes the fiber using polarization-multiplexed constant-power coded sequences that are suited for coexistence of DAS with WDM data transmission over the same fiber. We study the effect of both axisymmetric and anisotropic events on the two quantities. We demonstrate, through numerical simulations and lab experiments, the estimation of effective birefringence magnitude in static conditions, and the joint estimation of common phase and effective birefringence magnitude in the case of dynamic longitudinal strain and anisotropic transverse strain. This allows for event discrimination and increased sensitivity to disturbances that act transversely on the fiber, since polarization will be responsive to perturbations that break cylindrical symmetry, while the phase will strongly respond to longitudinal strain.
  • Robust brain age estimation from structural MRI with contrastive learning
    • Barbano Carlo Alberto
    • Dufumier Benoit
    • Duchesnay Edouard
    • Grangetto Marco
    • Gori Pietro
    Pattern Recognition Letters, Elsevier, 2026, 203, pp.78-84. <div><p>Estimating brain age from structural MRI has emerged as a powerful tool for characterizing normative and pathological aging. In this work, we explore contrastive learning as a scalable and robust alternative to L1-supervised approaches for brain age estimation. We introduce a novel contrastive loss function,  exp , and evaluate it across multiple public neuroimaging datasets comprising over 20,000 scans. Our experiments reveal four key findings. First, scaling pre-training on diverse, multi-site data consistently improves generalization performance, cutting external mean absolute error (MAE) nearly in half. Second,  exp is robust to site-related confounds, maintaining low scanner-predictability as training size increases. Third, contrastive models reliably capture accelerated aging in patients with cognitive impairment and Alzheimer's disease, as shown through brain age gap analysis, ROC curves, and longitudinal trends. Lastly, unlike L1-supervised baselines,  exp maintains a strong correlation between brain age accuracy and downstream diagnostic performance, supporting its potential as a foundation model for neuroimaging. These results position contrastive learning as a promising direction for building generalizable and clinically meaningful brain representations.</p></div> (10.1016/j.patrec.2026.02.032)
    DOI : 10.1016/j.patrec.2026.02.032
  • TaxoSurv: A Comprehensive Survey of Taxonomy Construction, Expansion, Completion, and Refinement
    • Ghamlouch Zeinab
    • Alam Mehwish
    , 2026. <div><p>Taxonomies are fundamental structures for organizing knowledge in the form of a hierarchy, supporting applications such as information retrieval, knowledge graphs, and semantic reasoning. However, many real-world taxonomies suffer from limited coverage, outdated concepts, and structural inconsistencies, motivating research on computational methods for constructing and refining hierarchical structures from heterogeneous data sources. This survey provides a systematic overview of taxonomy learning across its main tasks: construction, expansion, completion, and refinement. We introduce a structured categorization that organizes existing approaches along two dimensions, the nature of the downstream task and the methodology, including nonneural, neural, and LLM-based methods, and further analyze the benchmark datasets and evaluation protocols used across prior work. Our goal is to consolidate the state-of-the-art and propose a coherent taxonomy of methods that clarifies terminology and supports future research.</p></div>
  • Two-Indexed Schatten Quasi-Norms with Applications to Quantum Information Theory
    • Kochanowski Jan
    • Fawzi Omar
    • Rouzé Cambyse
    , 2026. We define 2-indexed $(q,p)$-Schatten quasi-norms for any $q,p &gt; 0$ on operators on a tensor product of Hilbert spaces, naturally extending the norms defined by Pisier's theory of operator-valued Schatten spaces. We establish several desirable properties of these quasi-norms, such as relational consistency and the behavior on block diagonal operators, assuming that $|\frac{1}{q} - \frac{1}{p}| \leq 1$. In fact, we show that this condition is essentially necessary for natural properties to hold. Furthermore, for linear maps between spaces of such quasi-norms, we introduce completely bounded quasi-norms and co-quasi-norms. We prove that the $q \to p$ completely bounded co-quasi-norm is super-multiplicative for tensor products of quantum channels for $q \geq p&gt;0$, extending an influential result of [Devetak, Junge, King, Ruskai, 2006]. Our proofs rely on elementary matrix analysis and operator convexity tools and do not require operator space theory. On the applications side, we demonstrate that these quasi-norms can be used to express relevant quantum information measures such as Rényi conditional entropies for $α\geq \frac{1}{2}$ or the Sandwiched Rényi Umlaut information for $α&lt; 1$. Our multiplicativity results imply a tensorizing notion of reverse hypercontractivity, additivity of the completely bounded minimum output Rényi-$α$-entropy for $α\geq\frac{1}{2}$ extending another important result of [Devetak, Junge, King, Ruskai, 2006], and additivity of the maximum output Rényi-$α$ entropy for $α\geq \frac{1}{2}$. (10.48550/arXiv.2604.14055)
    DOI : 10.48550/arXiv.2604.14055
  • FOLDER: Accelerating Multi-Modal Large Language Models with Enhanced Performance
    • Wang Haicheng
    • Yu Zhemeng
    • Spadaro Gabriele
    • Ju Chen
    • Quétu Victor
    • Xiao Shuai
    • Tartaglione Enzo
    , 2026, pp.23614-23625. Recently, Multi-modal Large Language Models (MLLMs) have shown remarkable effectiveness for multi-modal tasks due to their ability of cross-modal understanding. However, processing long sequences of visual tokens extracted from visual backbones poses challenges for deployment in real-time applications. To address this issue, we introduce FOLDER, a simple yet effective plug-and-play module designed to reduce the length of the visual token sequence, mitigating computational and memory demands during both training and inference. Through a comprehensive analysis of the token reduction process in the vision encoder, we analyze the information loss introduced by different reduction strategies and develop FOLDER to preserve key information while removing visual redundancy. We show the effectiveness of FOLDER by integrating it into the visual backbone of various MLLMs, significantly accelerating the inference phase. Furthermore, we evaluate its utility as a training accelerator or even performance booster for MLLMs. FOLDER achieves comparable or even better performance than the original models, while dramatically reducing complexity by removing up to 70 % of visual tokens. Our code is available at FOLDER Repository. (10.1109/ICCV51701.2025.02192)
    DOI : 10.1109/ICCV51701.2025.02192
  • Quantum Gibbs Sampling in Infinite Dimensions: Generation, Mixing Times and Circuit Implementation
    • Becker Simon
    • Rouzé Cambyse
    • Salzmann Robert
    , 2026. We develop a rigorous and implementable framework for Gibbs sampling of infinite-dimensional quantum systems governed by unbounded Hamiltonians. Extending dissipative Gibbs samplers beyond finite dimensions raises fundamental obstacles, including ill-defined generators, the absence of spectral gaps on natural Banach spaces, and tensions between implementability and convergence guarantees. We overcome these issues by constructing KMS-symmetric quantum Markov semigroups on separable Hilbert spaces that are both well-posed and efficiently implementable on qubit hardware. Our generation theory is based on the abstract framework of Dirichlet forms, adapted here to the case of algebras of bounded operators over separable Hilbert spaces. Leveraging the spectral properties of our self-adjoint generators, we establish quantitative convergence results in trace distance, including regimes of fast thermalization. In contrast, we also identify Hamiltonians for which a naive choice of generators guaranteeing implementability generally comes at the cost of losing convergence of the associated evolutions, thereby establishing a strong trade-off between implementability and convergence. Our framework applies to a wide class of models, including Schrödinger operators, Gaussian systems, and Bose-Hubbard Hamiltonians, and provides a unified approach linking rigorous infinite-dimensional analysis with algorithmic Gibbs state preparation. (10.48550/arXiv.2604.01192)
    DOI : 10.48550/arXiv.2604.01192
  • Generalised contextuality of continuous variable quantum theory can be revealed with a single projective measurement
    • Jokinen Pauli
    • Weilenmann Mirjam
    • Plávala Martin
    • Pellonpää Juha-Pekka
    • Kiukas Jukka
    • Uola Roope
    , 2026. Generalized contextuality is a possible indicator of non-classical behaviour in quantum information theory. In finite-dimensional systems, this is justified by the fact that noncontextual theories can be embedded into some simplex, i.e. into a classical theory. We show that a direct application of the standard definition of generalized contextuality to continuous variable systems does not envelope the statistics of some basic measurements, such as the position observable. In other words, we construct families of fully classical, i.e. commuting, measurements that nevertheless can be used to show contextuality of quantum theory. To overcome the apparent disagreement between the two notions of classicality, that is commutativity and noncontextuality, we propose a modified definition of generalised contextuality for continuous-variable systems. The modified definition is based on a physically-motivated approximation procedure, that uses only finite sets of measurement effects. We prove that in the limiting case this definition corresponds exactly to an extension of noncontextual models that benefits from non-constructive response functions. In the process, we discuss the extension of a known connection between contextuality and no-broadcasting to the continuous-variable scenario, and prove structural results regarding fixed points of infinite-dimensional entanglement breaking channels. (10.48550/arXiv.2601.14067)
    DOI : 10.48550/arXiv.2601.14067
  • Simulating Thermal Properties of Bose-Hubbard Models on a Quantum Computer
    • Becker Simon
    • Rouzé Cambyse
    • Salzmann Robert
    , 2026. While recent advances have established efficient quantum algorithms for preparing Gibbs states of finite-dimensional systems, comparable complexity results for bosonic and other infinite-dimensional models remain unexplored. We introduce the first general rigorous Gibbs sampling framework for bosonic many-body systems, showing that physically relevant bosonic models admit gapped dissipative generators, enabling efficient preparation of thermal states. Although our results hold for broad classes of models, we illustrate them using Bose-Hubbard Hamiltonians, both within and beyond the mean-field regime. In both cases, we show that the associated dissipative generators maintain a positive spectral gap, thereby implying exponential convergence to the thermal state. Our argument in the multi-mode case is based on a finite-rank reduction of the dissipative dynamics, which allows us to control the generator via compact perturbations and deduce the discreteness of the spectrum and the stability of the gap. We apply our results to provide efficient preparation of the corresponding Gibbs state on qubit hardware, and by that a quantum algorithm to compute thermal properties of the associated model. This provides the first mathematically controlled route to Gibbs sampling in infinite-dimensional systems, with implications for quantum simulation, thermalization, and many-body complexity, where quantum advantages may arise. (10.48550/arXiv.2604.06077)
    DOI : 10.48550/arXiv.2604.06077
  • Distortion-aware STAP adaptive beamforming for robust GNSS anti-jamming in high-dynamics environments
    • Lagarde Elise
    • Leborgne Fabien
    • Abedrrabba Sarra
    • Dubroca Norbert
    • Leonardon Mathieu
    • Roblin Christophe
    • Cousin Jean-Christophe
    • Gomes Joan
    • Oriol Stephane
    , 2026. he combination of Spatio-Temporal Adaptive Processing (STAP) and Controlled Reception Pattern Antennas (CRPAs) enhances the robustness of Global Navigation Satellite System (GNSS) anti-jamming receivers, particularly for autonomous launcher navigation where continuous Position, Velocity, and Time (PVT) availability is critical. In this architecture, STAP acts as a spatio-temporal filtering stage on array signals, while adaptive beamforming (ABF) algorithms compute and steer the filter weights to shape the antenna pattern and place spatial nulls under high-dynamics conditions involving strong carrier accelerations. Common techniques used to drive the STAP weights include Minimum Variance Distortionless Response (MVDR), Minimum Mean Square Error (MMSE), and Linearly Constrained Minimum Variance (LCMV). This paper demonstrates that, despite their theoretical optimality, such ABF-driven STAP approaches may struggle in complex real-world conditions, particularly in the presence of significant launcher dynamics, including rapid attitude variations and high acceleration profiles. To illustrate these limitations, two complementary sets of results are presented. The first is based on simulated data generated using open-source tools such as gnss-sdr-sim, allowing controlled emulation of dynamic interference and motion scenarios. The second relies on testbed experiments conducted in a controlled, conducted environment at the European Commission’s Joint Research Centre (JRC). The results show that variations in interference power, waveform type, and direction of arrival, combined with changes in GNSS signal geometry and carrier dynamics, introduce biases of varying magnitude. Significant nonlinear effects and distortions of the cross-correlation functions are observed, which degrade acquisition and tracking performance and may ultimately lead to PVT degradation or loss. These impairments are primarily induced by the conducted Radio Frequency (RF) chain, including front-end nonlinearities and gain increases introduced by the Low Noise Amplifier (LNA) and Automatic Gain Control (AGC), initially designed to improve GNSS signal detectability. Under such conditions, while adaptive beamforming algorithms may still compute STAP filter weights, the resulting interference cancellation may be insufficient to guarantee robust PVT tracking across all scenario configurations. More generally, classical adaptive beamforming approaches are sensitive to steering vector mismatch and covariance estimation errors induced by RF nonlinearities, high jammer-to-signal ratios (JSR), AGC saturation, and quantization effects. Compensating for these effects often requires increasing the number or complexity of constraints or regularization mechanisms, which in turn raises computational complexity, particularly in high-dynamics scenarios requiring frequent weight updates and highlights the need for accurate distortion modeling to properly dimension ABF constraints. This work highlights the need to revisit traditional ABF algorithms used to drive STAP filter weights to explicitly account for nonlinear and distortion effects, particularly in high-dynamics scenarios where robust PVT tracking is required. It also opens the way toward hybrid approaches combining classical digital signal processing with lightweight machine learning–based techniques, in which computationally frugal data-driven models assist in robust covariance estimation, distortion-aware constraint adaptation, and tracking of dynamic interference conditions, thereby improving robustness and maintaining reliable GNSS-based autonomy in RF contested environments.
  • Computing the free energy of quantum Coulomb gases and molecules via quantum Gibbs sampling
    • Becker Simon
    • Rouzé Cambyse
    • Salzmann Robert
    , 2026. We develop a quantum algorithm for estimating the free energy as well as the total Gibbs state of interacting quantum Coulomb gases and molecular systems in dimensions $d \in \{2,3\}$ at finite temperature. These systems lie beyond the reach of existing methods due to their singular interactions and infinite-dimensional Hilbert space structure. First, we show that the free energy of the full many-body Hamiltonian can be approximated by that of the same Hamiltonian with a finite-rank low-energy truncation of the interaction, with an explicit error bound polynomial in the particle number. This reduces the problem to a controlled finite-rank perturbation problem. Second, we introduce a quantum Gibbs sampling scheme tailored to this truncated system, based on a class of quantum Markov semigroups. Our main analytical result establishes that the associated generator has a strictly positive spectral gap for every truncation, implying exponential convergence to the target Gibbs state. This provides, to our knowledge, the first rigorous mixing-time guarantee for Gibbs sampling in a Coulomb interacting continuous-variable quantum system. Finally, we give an explicit quantum circuit implementation of the dynamics and derive an end-to-end complexity bound for approximating the free energy and the Gibbs state itself. Our results provide a mathematically rigorous route to quantum algorithms for free energy estimation in interacting quantum systems, without relying on classical approximations such as the Born-Oppenheimer reduction. (10.48550/arXiv.2604.15263)
    DOI : 10.48550/arXiv.2604.15263
  • A Low-Cost Multiplatform VLC System Prototype for Indoor Attocell Downlink Communication
    • Nascimento Alaf
    • Costa Wesley
    • Camporez Higor
    • Zwaag Klaas
    • Santos Francisco
    • Silva Jair
    , 2026. Devices operating in the Radio Frequency (RF) spectrum face capacity limitations due to licensed bandwidth and rapidly growing connectivity demand. Visible Light Communication (VLC) emerges as a complementary alternative, leveraging existing lighting infrastructure for energy-efficient indoor communication. However, for widespread adoption, VLC devices must achieve a cost comparable to established technologies. This work proposes a Plug-and-Play (PnP) USB dongle for VLC data reception based on low-cost hardware, along with a multiplatform interface for reactive data visualization. Our prototype achieves error-free downlink communication in attocell scenarios over distances up to 280 cm under line-of-sight (LOS) conditions.
  • Balanced Latent Semantics and Signal Fidelity for EEG Representation Learning
    • Nguyen Van-Chien
    • Tran Trung-Hieu
    • Doan Tuan-Kiet
    • Pham Quang Hung
    • Vu Ngoc-Son
    • Le Duc Han
    • Phan Huy
    • Le Nguyen Phi
    • Simidjievski Nikola
    • Tardieu Samuel
    • Nguyen Van-Tam
    , 2026. <div><p>Electroencephalography (EEG) is critical for neurological diagnosis but suffers from low SNR and subject variability. Current foundation models relying on raw signal reconstruction often overfit to local noise. We propose STELAR, a foundation model with a dual-space objective combining patch-level masked latent prediction for semantic stability with masked reconstruction for raw signal fidelity. To balance these objectives, we introduce MTPE-GB, a validation-driven gradient balancer that adaptively weights tasks without manual tuning or computational overhead. STELAR achieves state-of-the-art linear probing performance across diverse EEG benchmarks, demonstrating robust generalization.</p></div>
  • The ALERT Dataset: Benchmarking Anomaly Detection of Non-Stationary Vibrational Signals
    • Emelchenkov Anton
    • Fontaine Mathieu
    • Mahé Hervé
    • Roueff François
    , 2026. In recent years, automatic audio anomaly detection has gained considerable attention. However, most existing methods and benchmarks assume stationary or periodic signals, limiting their applicability to industrial environments characterized by non-stationary operating regimes such as speed ramps and transient load variations. We introduce the ALERT Dataset, a large-scale collection of non-stationary vibration recordings from electric powertrains acquired on an industrial end-of-line test bench. Each recording captures ramp-up and ramp-down phases with continuously varying rotational speed and includes synchronized speed measurements to enable explicit conditioning on operational dynamics. The dataset comprises 224 healthy training recordings and 80 healthy test recordings, along with an additional 80-sample hold-out set reserved for anomaly generation. From this hold-out set, multiple anomalous test suites (80 samples each) are constructed via expert-designed amplitude-based degradations and structured noise perturbations at varying signal-to-noise ratios, simulating realistic fault scenarios. Models are evaluated by discriminating these anomalies from the 80 healthy test recordings under a one-class learning paradigm. The benchmark further supports diverse protocols, including zero-shot cross-phase testing. To our knowledge, the ALERT Dataset is the first large-scale collection of non-stationary industrial vibration signals with synchronized speed references, addressing a critical gap in existing benchmarks. The dataset is publicly available on Zenodo. (10.5281/zenodo.18759681)
    DOI : 10.5281/zenodo.18759681
  • Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
    • Girrbach Leander
    • Alaniz Stephan
    • Smith Genevieve
    • Darrell Trevor
    • Akata Zeynep
    , 2026. Vision-language models trained on large-scale multimodal datasets show strong demographic biases, but the role of training data in producing these biases remains unclear. A major barrier has been the lack of demographic annotations in web-scale datasets such as LAION-400M. We address this gap by creating person-centric annotations for the full dataset, including over 276 million bounding boxes, perceived gender and race/ethnicity labels, and automatically generated captions. These annotations are produced through validated automatic labeling pipelines combining object detection, multimodal captioning, and finetuned classifiers. Using them, we uncover demographic imbalances and harmful associations, such as the disproportionate linking of men and individuals perceived as Black or Middle Eastern with crime-related and negative content. We also show that a linear fit predicts 60-70% of gender bias in CLIP and Stable Diffusion from direct co-occurrences in the data. Our resources establish the first large-scale empirical link between dataset composition and downstream model bias. Code is available at https://github.com/ExplainableML/LAION-400M-Person-Centric-Annotations.
  • NeuroSnitch: Exploiting Inter-Spike Interval Statistics for Timing Side-Channel Attacks on Noisy Neuromorphic Systems
    • Khan Mahreen
    • Mushtaq Maria
    • Apvrille Ludovic
    , 2026. <div><p>Neuromorphic computing promises energy-efficient solutions for embedded and edge systems, but introduces unique security challenges and a new attack surface. This paper presents NeuroSnitch, a first-ever timing side-channel attack to leverage subtle statistical variations in Inter-Spike Intervals (ISIs) on Spiking Neural Networks (SNNs) to extract secret information. We show that secret data, when modulating a neuron's input current, can be profiled through higher-order ISI statistics-mean, variance, skewness, and kurtosis-even under realistic noise sources, including observation noise, current fluctuation, and voltage jitter. Using the Leaky Integrateand-Fire (LIF) neuron model, we demonstrate that a Random Forest classifier can achieve 98.41% character-level classification accuracy on noisy ISI traces, enabling complete recovery of a 33-character secret string. This work exposes a previously underexplored and robust timing leakage vector in SNNs, underscoring the urgent need for tailored security measures in this emerging computing paradigm, particularly for sensitive embedded and IoT applications.</p></div> (10.1145/YYYYYYY.YYYYYYY)
    DOI : 10.1145/YYYYYYY.YYYYYYY
  • Study of Training Dynamics for Memory-Constrained Fine-Tuning
    • Quélennec Aël
    • Mozharovskyi Pavlo
    • Nguyen Van-Tam
    • Tartaglione Enzo
    , 2026. Memory-efficient training of deep neural networks has become increasingly important as models grow larger while deployment environments impose strict resource constraints. We propose TraDy, a novel transfer learning scheme leveraging two key insights: layer importance for updates is architecture-dependent and determinable a priori, while dynamic stochastic channel selection provides superior gradient approximation compared to static approaches. We introduce a dynamic channel selection approach that stochastically resamples channels between epochs within preselected layers. Extensive experiments demonstrate TraDy achieves state-of-the-art performance across various downstream tasks and architectures while maintaining strict memory constraints, achieving up to 99% activation sparsity, 95% weight derivative sparsity, and 97% reduction in FLOPs for weight derivative computation.