Sorry, you need to enable JavaScript to visit this website.
Share

Publications

2026

  • Conception d'un redresseur pour les bandes Wi-Fi et millimétrique
    • Wang Yibo
    • Niotaki Kyriaki
    • Lepage Anne Claire
    • Begaud Xavier
    , 2026.
  • [WIM AWARD] Nanogénérateurs Triboélectriques à Glissement Latéral pour la Récupération d’Énergie Hybride
    • Lu Yayi
    • Niotaki Kyriaki
    • Mohellebi Reda
    • Ben Kalaia Karim
    • Jabbour Chadi
    , 2026, pp.1. La récupération d'énergie RF est employée dans le but d'alimenter des appareils connectés à faible consommation. Pour augmenter et stabiliser le niveau de puissance récupéré à la sortie, la source RF est souvent associée avec d'autres types d'énergie dans un système complexe, on parle alors de récupération d'énergie hybride multi-sources. La triboélectrité basée sur le frottement a été repérée comme une voie émergente. Nos études paramétriques sur les nanogénérateurs triboélectriques à glissement latéral (Lateral Sliding-Triboelectric Nanogenerators/LS-TENGs) démontrent que : l'impédance de charge, la vitesse de déplacement et la surface de contact sont parmi les principaux facteurs à considérer, afin de maximiser la puissance récupérée.
  • Energy efficiency of quantum computers
    • Carrasco-Codina Miquel
    • Escofet Pau
    • Hilaire Paul
    • Soret Ariane
    • Nerenberg Sam
    • Champain Victor
    • Milburn Gerard
    • Theophilo Klara
    • Li Sophie
    • Bautista Irais
    • Gómez Andrés
    • Miralles Jose
    • Abadal Sergi
    • Almudéver Carmen
    • Alarcón Eduard
    • Yehia Raja
    , 2026. How much energy does a quantum computer consume? Are they more efficient than their classical counterparts? In this work, we make a step towards answering these questions. We define the energy efficiency of a quantum computer as the ratio of the number of algorithms it can perform during a given time over the energy consumed by the hardware during this time. We analyze the most representative physical platforms currently envisioned to be used as building blocks of quantum computers: superconducting qubits, silicon spin qubits, trapped ions, neutral atoms and photonic qubits. Including insights from experts in all these technologies and taking into account algorithm compilation constraints, we discuss the advantages and inconveniences of each platform from an energy standpoint. Beyond providing concrete values of the energy consumption of current quantum computers, we lay the foundation of a framework to benchmark the energy efficiency of any future quantum computing architecture. (10.48550/arXiv.2605.15090)
    DOI : 10.48550/arXiv.2605.15090
  • Complexity of unicity problems
    • Hudry Olivier
    , 2026.
  • SPOT: An Annotated French Corpus and Benchmark for Detecting Critical Interventions in Online Conversations
    • Berriche Manon
    • Nouri Célia
    • Clavel Chloé
    • Cointet Jean-Philippe
    , 2026. We introduce SPOT (Stopping Points in Online Threads), the first annotated corpus translating the sociological concept of stopping point into a reproducible NLP task. Stopping points are ordinary critical interventions that pause or redirect online discussions through a range of forms (irony, subtle doubt or fragmentary arguments) that frameworks like counterspeech or social correction often overlook. We operationalize this concept as a binary classification task and provide reliable annotation guidelines. The corpus contains 43,305 manually annotated French Facebook comments linked to URLs flagged as false information by social media users, enriched with contextual metadata (article, post, parent comment, page or group, and source). We benchmark fine-tuned encoder models (CamemBERT) and instruction-tuned LLMs under various prompting strategies. Results show that fine-tuned encoders outperform prompted LLMs in F1 score by more than 10 percentage points, confirming the importance of supervised learning for emerging non-English social media tasks. Incorporating contextual metadata further improves encoder models F1 scores from 0.75 to 0.78. We release the anonymized dataset, along with the annotation guidelines and code in our code repository, to foster transparency and reproducible research.
  • Experimental Evaluation of QoT Uncertainty Reduction Techniques in Network Digital Twins
    • Purkayastha Ambashri
    • Delezoide Camille
    • Boitier Fabien
    • Lourdiane Mounia
    • Ware Cédric
    • Layec Patricia
    , 2026. <div><p>We experimentally demonstrate sequential parameter refinement leveraging lightpath diversity in presence of various levels of uncertainty on parameter knowledge to iteratively improve QoT prediction accuracy by up to 1.8 dB.</p></div>
  • ENEIDE: A High Quality Silver Standard Dataset for Named Entity Recognition and Linking in Historical Italian
    • Santini Cristian
    • Barzaghi Sebastian
    • Sernani Paolo
    • Frontoni Emanuele
    • Melosi Laura
    • Alam Mehwish
    , 2026. This paper introduces ENEIDE (Extracting Named Entities from Italian Digital Editions), a silver standard dataset for Named Entity Recognition and Linking (NERL) in historical Italian texts. The corpus comprises 2,111 documents with over 8,000 entity annotations semi-automatically extracted from two scholarly digital editions: Digital Zibaldone, the philosophical diary of the Italian poet Giacomo Leopardi (1798-1837), and Aldo Moro Digitale, the complete works of the Italian politician Aldo Moro (1916-1978). Annotations cover multiple entity types (person, location, organization, literary work) linked to Wikidata identifiers, including NIL entities that cannot be mapped to the knowledge graph. To the best of our knowledge, ENEIDE represents the first multi-domain, publicly available NERL dataset for historical Italian with training, development, and test splits. We present a methodology for semi-automatic annotations extraction from manually curated scholarly digital editions, including quality control and annotation enhancement procedures. Baseline experiments using state-of-the-art models demonstrate the dataset's challenge for NERL and the gap between zero-shot approaches and fine-tuned models. The dataset's diachronic coverage spanning two centuries makes it particularly suitable for temporal entity disambiguation and cross-domain evaluation.
  • Analysing Lightweight Large Language Models for Biomedical Named Entity Recognition on Diverse Ouput Formats
    • Epron Pierre
    • Coulet Adrien
    • Alam Mehwish
    , 2026, pp.1-13. Despite their strong linguistic capabilities, Large Language Models (LLMs) are computationally demanding and require substantial resources for fine-tuning, which is unadapted to privacy and budget constraints of many healthcare settings. To address this, we present an experimental analysis focused on Biomedical Named Entity Recognition using lightweight LLMs, we evaluate the impact of different output formats on model performance. The results reveal that lightweight LLMs can achieve competitive performance compared to the larger models, highlighting their potential as lightweight yet effective alternatives for biomedical information extraction. Our analysis shows that instruction tuning over many distinct formats does not improve performance, but identifies several format consistently associated with better performance.
  • POT Python Optimal Transport
    • Flamary Rémi
    • Vincent-Cuaz Cédric
    • Courty Nicolas
    • Gramfort Alexandre
    • Kachaiev Oleksii
    • Quang Tran Huy
    • David Laurène
    • Bonet Clément
    • Cassereau Nathan
    • Gnassounou Theo
    • Tanguy Eloi
    • Delon Julie
    • Collas Antoine
    • Mazelet Sonia
    • Chapel Laetitia
    • Kerdoncuff Tanguy
    • Yu Xizheng
    • Feickert Matthew
    • Krzakala Paul
    • Liu Tianlin
    • Fernandes Montesuma Eduardo
    , 2026. (10.5281/ZENODO.17161062)
    DOI : 10.5281/ZENODO.17161062
  • Leveraging machine learning for efficient and secure optical networking
    • Andrenacci Isaia
    , 2026. Future optical networks face rapidly growing traffic demands while operating close to fundamental capacity limits, making efficient network resource utilization increasingly critical. Massive monitoring and telemetry have emerged as a possible solution for more autonomous and adaptive network operation. In this PhD, I investigate machine learning-based techniques that leverage receiver-side measurements and optical network parameters to enhance the efficiency, reliability, and security of point-to-point optical fiber links. First, I propose machine learning models to accurately predict nonlinear interference spectrum components in wavelength-division multiplexed systems, enabling low-complexity and high-accuracy performance estimation. I then develop closed-loop machine learning controllers to optimize launch power in links affected by polarization-dependent loss using only receiver-side signal-to-noise ratio statistics. Furthermore, I introduce low-cost monitoring techniques to localize anomalous polarization-dependent loss elements and to identify the dominant noise in the link, facilitating future adaptive optimization and self-healing of the network. Finally, I propose a method to quantify fiber macro-bending events from spectral measurements, allowing future discrimination between maintenance activities and physical-layer attacks. Overall, this work demonstrates that practical machine learning methods, combined with pervasive monitoring, can support near-zero-margin operation and pave the way toward intelligent, autonomous, and secure next-generation optical networks.
  • Kalman Filtering for Sensing Aided Communication to Mobile Users in Large Cellular Networks
    • Balakrishnan Ashutosh
    • Soprano-Loto Nahuel
    • Baccelli François
    , 2026. Upcoming 6G networks are expected to have joint sensing and communication capabilities. In this work, we propose a sensing-aided communication (SAC) framework, wherein a cellular network of randomly located base stations (BSs) sequentially estimates the position and channel gain of a user equipment (UE) moving randomly across the network. At the core of this framework lies a state-dependent Kalman filter that extends the classical formulation in two key respects: (i) measurements are acquired from a dynamically selected BS, the selection based on the a-priori estimate of the UE's position and fading channel; and (ii) the measurement noise covariance is modelled as a function of the sensing distance and fading, coupling the temporal evolution of the state with the measurement quality. This coupling gives rise to pronounced estimation error peaks during handovers, which in turn directly govern the design of initial access and handover protocols in stochastic networks. To quantify performance, we introduce a new metric, the stationary mean of the estimation error covariance, and establish its validity through an ergodic theorem, further supported by numerical convergence studies. The sensing-based UE-BS association is evaluated in one-and two-dimensional Gauss-Markov mobility scenarios and benchmarked against two alternatives: an autonomous noise covariance setting and an ideal omniscient controller with perfect position knowledge. The SAC framework is shown to outperform a position-only sensing framework by around 32%. Overall, the results highlight the advantage of the SAC framework in providing protocol-level insights for handover-aware sensing and demonstrate that the proposed filter achieves stable long-run performance.
  • Actualizing Ethical Principles for Curating Large-Scale Training Datasets in the Era of Massive AI Models
    • Cazacu Silvia
    • Qian Alice
    • Zhao Dora
    • Pine Kathleen
    • Walker Shawn
    • Shen Hong
    • Dabbish Laura
    • Panagiotidou Georgia
    • Klaus Scheuerman Morgan
    , 2026, 21. While AI technologies are often framed as ubiquitous and inevitable, their expansion relies on the large-scale extraction of data from diverse global communities. However, the datasets powering foundation models are often treated as found artifacts rather than products of specific power dynamics and human labor. Current practices frequently disregard the structural inequities embedded in data, even as these systems profoundly impact systemically marginalized communities. While frameworks for ethical curation exist for smaller datasets, the massive scale of foundation models has introduced a logic of extraction that prioritizes volume over accountability. This workshop invites researchers, practitioners, and activists to move beyond standard technical hurdles and instead problematize the foundational assumptions of large-scale data work. This workshop builds on a series of ongoing conversations across scholarly communities and continues work initiated at CSCW 2025. Drawing from the CRAFT tradition of transdisciplinary exchange and community action, we will facilitate a collective refactoring of the three core pillars of data curation: composition, focused on interrogating whose lives are extracted and how representation is shaped by hegemonic interests; process, which centers the invisible labor and situated contexts involved in curating and cleaning massive data stores, and release, focused on rethinking the governance, accountability, and potential for refusal in how these models are shared with the world. Our goal is to cultivate a community-led conceptual framework that reimagines more just sociotechnical futures—transforming data curation from a top-down technical requirement into an act of collective responsibility and repair. (10.1108/jices-08-2022-0069)
    DOI : 10.1108/jices-08-2022-0069
  • RAVE: Rate-Adaptive Visual Encoding for 3D Gaussian Splatting
    • Tran Hoang-Nhat
    • Di Sario Francesco
    • Spadaro Gabriele
    • Valenzise Giuseppe
    • Tartaglione Enzo
    , 2026, pp.11727-11731. <div><p>Recent advances in neural scene representations have transformed immersive multimedia, with 3D Gaussian Splatting (3DGS) enabling real-time photorealistic rendering. Despite its efficiency, 3DGS suffers from large memory requirements and costly training procedures, motivating efforts toward compression. Existing approaches, however, operate at fixed rates, limiting adaptability to varying bandwidth and device constraints. In this work, we propose a flexible compression scheme for 3DGS that supports interpolation at any rate between predefined bounds. Our method is computationally lightweight, requires no retraining for any rate, and preserves rendering quality across a broad range of operating points. Experiments demonstrate that the approach achieves efficient, high-quality compression while offering dynamic rate control, making it suitable for practical deployment in immersive applications. The code is available at https://github.com/inspiros/RAVE.</p></div> (10.1109/ICASSP55912.2026.11463333)
    DOI : 10.1109/ICASSP55912.2026.11463333
  • S-PRESSO: Ultra Low Bitrate Sound Effect Compression With Diffusion Autoencoders And Offline Quantization
    • Lahrichi Zineb
    • Hadjeres Gaëtan
    • Richard Gaël
    • Peeters Geoffroy
    , 2026. <div><p>Neural audio compression models have recently achieved extreme compression rates, enabling efficient latent generative modeling. Conversely, latent generative models have been applied to compression, pushing the limits of continuous and discrete approaches. However, existing methods remain constrained to low-resolution audio and degrade substantially at very low bitrates, where audible artifacts are prominent. In this paper, we present S-PRESSO, a 48kHz sound effect compression model that produces both continuous and discrete embeddings at ultra-low bitrates, down to 0.096 kbps, via offline quantization. Our model relies on a pretrained latent diffusion model to decode compressed audio embeddings learned by a latent encoder. Leveraging the generative priors of the diffusion decoder, we achieve extremely low frame rates, down to 1Hz (750x compression rate), producing convincing and realistic reconstructions at the cost of exact fidelity. Despite operating at high compression rates, we demonstrate that S-PRESSO outperforms both continuous and discrete baselines in audio quality, acoustic similarity and reconstruction metrics.</p></div>
  • SIRUP: A DIFFUSION-BASED VIRTUAL UPMIXER OF STEERING VECTORS FOR HIGHLY-DIRECTIVE SPATIALIZATION WITH FIRST-ORDER AMBISONICS
    • Picard Emilio
    • Carlo Diego Di
    • Nugraha Aditya Arie
    • Fontaine Mathieu
    • Yoshii Kazuyoshi
    , 2026, pp.14707-14711. <div><p>This paper presents virtual upmixing of steering vectors captured by a fewer-channel spherical microphone array. This challenge has conventionally been addressed by recovering the directions and signals of sound sources from first-order ambisonics (FOA) data, and then rendering the higher-order ambisonics (HOA) data using a physics-based acoustic simulator. This approach, however, struggles to handle the mutual dependency between the spatial directivity of source estimation and the spatial resolution of FOA ambisonics data. Our method, named SIRUP, employs a latent diffusion model architecture. Specifically, a variational autoencoder (VAE) is used to learn a compact encoding of the HOA data in a latent space and a diffusion model is then trained to generate the HOA embeddings, conditioned by the FOA data. Experimental results showed that SIRUP achieved a significant improvement compared to FOA systems for steering vector upmixing, source localization, and speech denoising.</p></div> (10.1109/ICASSP55912.2026.11464234)
    DOI : 10.1109/ICASSP55912.2026.11464234
  • HFMCA: Orthonormal Feature Learning for EEG-Based Brain Decoding
    • Wang Yinghao
    • Xu Lintao
    • Yu Shujian
    • Tartaglione Enzo
    • Nguyen Van-Tam
    , 2026, pp.6826-6830. Electroencephalography (EEG) analysis is critical for brain-computer interfaces and neuroscience, but the intrinsic noise and high dimensionality of EEG signals hinder effective feature learning. We propose a self-supervised framework based on the Hierarchical Functional Maximal Correlation Algorithm (HFMCA), which learns orthonormal EEG representations by enforcing feature decorrelation and reducing redundancy. This design enables robust capture of essential brain dynamics for various EEG recognition tasks. We validate HFMCA on two benchmark datasets, SEED and BCIC-2A, where pretraining with HFMCA consistently outperforms competitive self-supervised baselines, achieving notable gains in classification accuracy. Across diverse EEG tasks, our method demonstrates superior cross-subject generalization under leave-one-subject-out validation, advancing state-of-the-art by 2.71% on SEED emotion recognition and 2.57% on BCIC-2A motor imagery classification. Our code and supplementary material are available at: https://github.com/W-Yinghao/HFMCA_EEG. (10.1109/ICASSP55912.2026.11464747)
    DOI : 10.1109/ICASSP55912.2026.11464747
  • Residual Tokens Enhance Masked Autoencoders For Speech Modeling
    • Sadok Samir
    • Lathuilière Stéphane
    • Alameda-Pineda Xavier
    , 2026, pp.14447-14451. Recent speech modeling relies on explicit attributes such as pitch, content, and speaker identity, but these alone cannot capture the full richness of natural speech. We introduce RT-MAE, a novel masked autoencoder framework that augments the supervised attributes-based modeling with unsupervised residual trainable tokens, designed to encode the information not explained by explicit labeled factors (e.g., timbre variations, noise, emotion etc). Experiments show that RT-MAE improves reconstruction quality, preserving content and speaker similarity while enhancing expressivity. We further demonstrate its applicability to speech enhancement, removing noise at inference while maintaining controllability and naturalness. (10.1109/ICASSP55912.2026.11464180)
    DOI : 10.1109/ICASSP55912.2026.11464180
  • PHYSICS-INFORMED LEARNING OF NEURAL SCATTERING FIELDS TOWARDS MEASUREMENT-FREE MESH-TO-HRTF ESTIMATION
    • Martinez Tancrède
    • Carlo Diego Di
    • Nugraha Aditya Arie
    • Fontaine Mathieu
    • Yoshii Kazuyoshi
    , 2026, pp.22577-22581. <div><p>This paper describes neural simulation of the scattered pressure field from a plane wave around a scattering object in both continuous 2D and 3D domains. This task has typically been treated as a regression problem that aims to train a physicsinformed neural network (PINN) using pressure measurements at discrete positions. This approach, however, needs to train the whole network for each incident wave direction. To address this, we propose a measurement-free simulator based on a PINN purely driven by the Helmholtz equation with the Robin boundary condition and the Sommerfeld radiation condition with the aid of the perfectly matched layer (PML) framework. More specifically, we design a physics-informed scattering hypernetwork (PHISK) that can generalize to incident waves from any direction via low-rank adaptation (LoRA) of a PINN trained for a specific configuration. The experiment shows that the proposed method accurately simulated sound scattering around various objects, adapting to unseen incident wave directions with minimal performance loss, and realized reasonable simulation of head-related transfer functions (HRTFs) from complex mesh data of a human head.</p></div> (10.1109/ICASSP55912.2026.11462698)
    DOI : 10.1109/ICASSP55912.2026.11462698
  • Generalization Bounds for Spectral GNNs via Fourier Domain Analysis
    • Martirosyan Vahan A
    • Malitesta Daniele
    • Talbot Hugues
    • Giraldo Jhony H
    • Malliaros Fragkiskos D
    , 2026, 300. Spectral graph neural networks learn graph filters, but their behavior with increasing depth and polynomial order is not well understood. We analyze these models in the graph Fourier domain, where each layer becomes an element-wise frequency update, separating the fixed spectrum from trainable parameters and making depth and order explicit. In this setting, we show that Gaussian complexity is invariant under the Graph Fourier Transform, which allows us to derive data-dependent, depth, and order-aware generalization bounds together with stability estimates. In the linear case, our bounds are tighter, and on real graphs, the data-dependent term correlates with the generalization gap across polynomial bases, highlighting practical choices that avoid frequency amplification across layers.
  • Numerical Model of a Multiple-Input-Multiple-Output Distributed Acoustic Sensor with Joint Phase and Birefringence Estimation
    • Prato Diane
    • Sheramin Mehran Mokhtari
    • Gabet Renaud
    • Awwad Élie
    , 2026. In this work, we introduce and experimentally validate a numerical model for a Multiple-Input-Multiple-Output Distributed Acoustic Sensing (MIMO-DAS) system that accounts for dynamic perturbations of fiber birefringence and of the common optical phase of the backscattered signal (or polarization-averaged phase, shared by both polarization tributaries). The MIMO-DAS system probes the fiber using polarization-multiplexed constant-power coded sequences that are suited for coexistence of DAS with WDM data transmission over the same fiber. We study the effect of both axisymmetric and anisotropic events on the two quantities. We demonstrate, through numerical simulations and lab experiments, the estimation of effective birefringence magnitude in static conditions, and the joint estimation of common phase and effective birefringence magnitude in the case of dynamic longitudinal strain and anisotropic transverse strain. This allows for event discrimination and increased sensitivity to disturbances that act transversely on the fiber, since polarization will be responsive to perturbations that break cylindrical symmetry, while the phase will strongly respond to longitudinal strain.
  • Robust brain age estimation from structural MRI with contrastive learning
    • Barbano Carlo Alberto
    • Dufumier Benoit
    • Duchesnay Edouard
    • Grangetto Marco
    • Gori Pietro
    Pattern Recognition Letters, Elsevier, 2026, 203, pp.78-84. <div><p>Estimating brain age from structural MRI has emerged as a powerful tool for characterizing normative and pathological aging. In this work, we explore contrastive learning as a scalable and robust alternative to L1-supervised approaches for brain age estimation. We introduce a novel contrastive loss function,  exp , and evaluate it across multiple public neuroimaging datasets comprising over 20,000 scans. Our experiments reveal four key findings. First, scaling pre-training on diverse, multi-site data consistently improves generalization performance, cutting external mean absolute error (MAE) nearly in half. Second,  exp is robust to site-related confounds, maintaining low scanner-predictability as training size increases. Third, contrastive models reliably capture accelerated aging in patients with cognitive impairment and Alzheimer's disease, as shown through brain age gap analysis, ROC curves, and longitudinal trends. Lastly, unlike L1-supervised baselines,  exp maintains a strong correlation between brain age accuracy and downstream diagnostic performance, supporting its potential as a foundation model for neuroimaging. These results position contrastive learning as a promising direction for building generalizable and clinically meaningful brain representations.</p></div> (10.1016/j.patrec.2026.02.032)
    DOI : 10.1016/j.patrec.2026.02.032
  • TaxoSurv: A Comprehensive Survey of Taxonomy Construction, Expansion, Completion, and Refinement
    • Ghamlouch Zeinab
    • Alam Mehwish
    , 2026. <div><p>Taxonomies are fundamental structures for organizing knowledge in the form of a hierarchy, supporting applications such as information retrieval, knowledge graphs, and semantic reasoning. However, many real-world taxonomies suffer from limited coverage, outdated concepts, and structural inconsistencies, motivating research on computational methods for constructing and refining hierarchical structures from heterogeneous data sources. This survey provides a systematic overview of taxonomy learning across its main tasks: construction, expansion, completion, and refinement. We introduce a structured categorization that organizes existing approaches along two dimensions, the nature of the downstream task and the methodology, including nonneural, neural, and LLM-based methods, and further analyze the benchmark datasets and evaluation protocols used across prior work. Our goal is to consolidate the state-of-the-art and propose a coherent taxonomy of methods that clarifies terminology and supports future research.</p></div>
  • FlowC2S: Flowing from Current to Succeeding Frames for Fast and Memory-Efficient Video Continuation
    • Margaryan Hovhannes
    • Bammey Quentin
    • Sandor Christian
    , 2026. This paper introduces a novel methodology for generating fast and memory-efficient video continuations. Our method, dubbed FlowC2S, fine-tunes a pre-trained text-to-video flow model to learn a vector field between the current and succeeding video chunks. Two design choices are key. First, we introduce inherent optimal couplings, utilizing temporally adjacent video chunks during training as a practical proxy for true optimal couplings, resulting in straighter flows. Second, we incorporate target inversion, injecting the inverted latent of the target chunk into the input representation to strengthen correspondences and improve visual fidelity. By flowing directly from current to succeeding frames, instead of the common combination of current frames with noise to generate a video continuation, we reduce the dimensionality of the model input by a factor of two. The proposed method, fine-tuned from LTXV and Wan, surpasses the state-of-the-art scores across quantitative evaluations with FID and FVD, with as few as five neural function evaluations.
  • Two-Indexed Schatten Quasi-Norms with Applications to Quantum Information Theory
    • Kochanowski Jan
    • Fawzi Omar
    • Rouzé Cambyse
    , 2026. We define 2-indexed $(q,p)$-Schatten quasi-norms for any $q,p &gt; 0$ on operators on a tensor product of Hilbert spaces, naturally extending the norms defined by Pisier's theory of operator-valued Schatten spaces. We establish several desirable properties of these quasi-norms, such as relational consistency and the behavior on block diagonal operators, assuming that $|\frac{1}{q} - \frac{1}{p}| \leq 1$. In fact, we show that this condition is essentially necessary for natural properties to hold. Furthermore, for linear maps between spaces of such quasi-norms, we introduce completely bounded quasi-norms and co-quasi-norms. We prove that the $q \to p$ completely bounded co-quasi-norm is super-multiplicative for tensor products of quantum channels for $q \geq p&gt;0$, extending an influential result of [Devetak, Junge, King, Ruskai, 2006]. Our proofs rely on elementary matrix analysis and operator convexity tools and do not require operator space theory. On the applications side, we demonstrate that these quasi-norms can be used to express relevant quantum information measures such as Rényi conditional entropies for $α\geq \frac{1}{2}$ or the Sandwiched Rényi Umlaut information for $α&lt; 1$. Our multiplicativity results imply a tensorizing notion of reverse hypercontractivity, additivity of the completely bounded minimum output Rényi-$α$-entropy for $α\geq\frac{1}{2}$ extending another important result of [Devetak, Junge, King, Ruskai, 2006], and additivity of the maximum output Rényi-$α$ entropy for $α\geq \frac{1}{2}$. (10.48550/arXiv.2604.14055)
    DOI : 10.48550/arXiv.2604.14055
  • FOLDER: Accelerating Multi-Modal Large Language Models with Enhanced Performance
    • Wang Haicheng
    • Yu Zhemeng
    • Spadaro Gabriele
    • Ju Chen
    • Quétu Victor
    • Xiao Shuai
    • Tartaglione Enzo
    , 2026, pp.23614-23625. Recently, Multi-modal Large Language Models (MLLMs) have shown remarkable effectiveness for multi-modal tasks due to their ability of cross-modal understanding. However, processing long sequences of visual tokens extracted from visual backbones poses challenges for deployment in real-time applications. To address this issue, we introduce FOLDER, a simple yet effective plug-and-play module designed to reduce the length of the visual token sequence, mitigating computational and memory demands during both training and inference. Through a comprehensive analysis of the token reduction process in the vision encoder, we analyze the information loss introduced by different reduction strategies and develop FOLDER to preserve key information while removing visual redundancy. We show the effectiveness of FOLDER by integrating it into the visual backbone of various MLLMs, significantly accelerating the inference phase. Furthermore, we evaluate its utility as a training accelerator or even performance booster for MLLMs. FOLDER achieves comparable or even better performance than the original models, while dramatically reducing complexity by removing up to 70 % of visual tokens. Our code is available at FOLDER Repository. (10.1109/ICCV51701.2025.02192)
    DOI : 10.1109/ICCV51701.2025.02192