Sorry, you need to enable JavaScript to visit this website.
Partager

Publications

 

Les publications de nos enseignants-chercheurs sont sur la plateforme HAL :

 

Les publications des thèses des docteurs du LTCI sont sur la plateforme HAL :

 

Retrouver les publications figurant dans l'archive ouverte HAL par année :

2023

  • Tailored vertex ordering for faster triangle listing in large graphs
    • Lécuyer Fabrice
    • Jachiet Louis
    • Magnien Clémence
    • Tabourier Lionel
    , 2023. Listing triangles is a fundamental graph problem with many applications, and large graphs require fast algorithms. Vertex ordering allows the orientation of edges from lower to higher vertex indices, and state-of-the-art triangle listing algorithms use this to accelerate their execution and to bound their time complexity. Yet, only basic orderings have been tested. In this paper, we show that studying the precise cost of algorithms instead of their bounded complexity leads to faster solutions. We introduce cost functions that link ordering properties with the running time of a given algorithm. We prove that their minimization is NP-hard and propose heuristics to obtain new orderings with different trade-offs between cost reduction and ordering time. Using datasets with up to two billion edges, we show that our heuristics accelerate the listing of triangles by an average of 38% when the ordering is already given as an input, and 16% when the ordering time is included. (10.1137/1.9781611977561.ch7)
    DOI : 10.1137/1.9781611977561.ch7
  • Reconfigurable Adaptive Channel Sensing
    • Mukherjee Manuj
    • Tchamkerten Aslan
    • Jabbour Chadi
    IEEE Transactions on Green Communications and Networking, IEEE, 2023, 7 (3), pp.1394 - 1406. This paper proposes an energy-efficient detection scheme, referred to as AdaSense, that is particularly suitable in the sparse regime when events to be detected happen rarely. To minimize energy consumption, AdaSense exploits the dependency between the receiver noise figure (i.e., the receiver added noise) and the receiver power consumption; less noisy channel observations typically imply higher power consumption. AdaSense is duty-cycled and begins each cycle with a few channel observations in a low-power-low-reliability mode. Based on these observations, it makes a first tentative decision on whether or not a message is present. If no message is declared, AdaSense waits till the beginning of the next cycle and starts afresh. If a message is tentatively declared, AdaSense enters a confirmation second phase, takes more samples, but now in a high-power-high-reliability mode. If these observations confirm the tentative decision, AdaSense stops, else AdaSense waits till the beginning of the next cycle and starts afresh in the low-power-low-reliability mode. Compared to prominent detection schemes such as the clear channel assessment algorithm of the Berkeley Media Access Control (BMAC) protocol, AdaSense provides relative energy gains that grow unbounded in the small probability of false-alarm regime, as communication gets sparser. In the non-asymptotic regime, energy gains are 30% to 75% for communication scenarios typically found in the context of wake-up receivers. (10.1109/TGCN.2023.3238176)
    DOI : 10.1109/TGCN.2023.3238176
  • Multi-temporal speckle reduction with self-supervised deep neural networks
    • Meraoumia Inès
    • Dalsasso Emanuele
    • Denis Loïc
    • Abergel Rémy
    • Tupin Florence
    IEEE Transactions on Geoscience and Remote Sensing, Institute of Electrical and Electronics Engineers, 2023, 61. Speckle filtering is generally a prerequisite to the analysis of synthetic aperture radar (SAR) images. Tremendous progress has been achieved in the domain of single-image despeckling. Latest techniques rely on deep neural networks to restore the various structures and textures peculiar to SAR images. The availability of time series of SAR images offers the possibility of improving speckle filtering by combining different speckle realizations over the same area. The supervised training of deep neural networks requires ground-truth speckle-free images. Such images can only be obtained indirectly through some form of averaging, by spatial or temporal integration, and are imperfect. Given the potential of very high quality restoration reachable by multi-temporal speckle filtering, the limitations of ground-truth images need to be circumvented. We extend a recent self-supervised training strategy for single-look complex SAR images, called MERLIN, to the case of multi-temporal filtering. This requires modeling the sources of statistical dependencies in the spatial and temporal dimensions as well as between the real and imaginary components of the complex amplitudes. Quantitative analysis on datasets with simulated speckle indicates a clear improvement of speckle reduction when additional SAR images are included. Our method is then applied to stacks of TerraSAR-X images and shown to outperform competing multitemporal speckle filtering approaches. The code of the trained models and supplementary results are made freely available at https://gitlab.telecom-paris.fr/ring/ multi-temporal-merlin/. (10.1109/TGRS.2023.3237466)
    DOI : 10.1109/TGRS.2023.3237466
  • Design Optimization of 12-Core Amplifier based on Erbium Ytterbium Co-doped Fiber for Spatial Multiplexed Transmission System
    • Lebreton Aurelien
    • Melin Gilles
    • Bordais Sylvain
    • Kerampran Romain
    • Pincemin Erwan
    • Taunay Thierry
    • Jauffrit Jeremie
    • Disez Pierre-Yves
    • Le Bouette Claude
    • Jaouen Yves
    • Morvan Michel
    • Lu Chao
    Journal of Lightwave Technology, Institute of Electrical and Electronics Engineers (IEEE)/Optical Society of America(OSA), 2023, 41 (02), pp.462 - 476. A 20 dB gain 12 cores Er 3+ /Yb 3+ co-doped cladding pumped amplifier in C-band with only 5.3 W of pump power has been achieved. A classical rate equation model has been applied for the amplifier design. Parameters such as active fiber length, pump power and ions concentration have been investigated and optimized. Results obtained through numerical simulation and experimental investigations are compared. Different use cases of MC-EYDFA have been studied, such as various transmission configuration or multi-core amplifier in ROADM architectures. 1200-km with 200G DP-QPSK and 300 km with 400G DP-16QAM are achieved in serial configuration at 1550 nm. This is a first step towards SDM transmission using power efficient amplifiers, for cost, energy and footprint saving. (10.1109/JLT.2022.3217308)
    DOI : 10.1109/JLT.2022.3217308
  • Iiro Honkala’s contributions to some conjectures on identifying codes
    • Hudry Olivier
    , 2023.
  • AI models for digital signal processing in future 6G-IoT networks
    • Larue Guillaume
    , 2023. Wireless technologies are of paramount importance to today's societies and future 6th generation communication networks are expected to address many societal and technological challenges. While communications infrastructures have a growing environmental impact that needs to be reduced, digital technologies also have a role to play in reducing the impact of all sectors of the economy. To this end, the future networks will not only have to enable more efficient information transfer, but also meet the growing need for data exchange capacity. This is particularly the role of the Internet of Things use cases, where a massive number of sensors allow to monitor complex systems. These use cases are associated with many constraints such as limited energy resources and complexity. Therefore, an efficient and low-complexity physical layer - responsible for the transmission of information between the network nodes - is absolutely crucial. In this regard, the use of artificial intelligence techniques is relevant. On the one hand, the mathematical framework of neural networks allows for efficient and low-cost generic hardware implementations. On the other hand, the application of learning procedures can improve the performance of certain algorithms. In this work, we are interested in the use of neural networks and machine learning for digital signal processing in the context of 6G-IoT networks. First, we are interested in the transcription of certain equalisation, demodulation and decoding algorithms from the digital communications literature into neural networks. Secondly, we are interested in the application of learning mechanisms on these neural network structures in order to improve their performance. A linear block decoder is proposed which allows the blind discovery of a decoding scheme whose performance is at least equivalent to that of the reference decoder. Finally, an end-to-end structure is presented, allowing joint learning of an encoding/decoding scheme with performance and complexity comparable to state-of-the-art solutions. (10.70675/b70858e5z6800z430dza632z33528def55bb)
    DOI : 10.70675/b70858e5z6800z430dza632z33528def55bb
  • Function-valued regression with kernels : Improving speed, flexibility and robustness
    • Bouche Dimitri
    , 2023. With the increasing ubiquity of data-collecting devices, a great variety of phenomena is monitored with finer and finer accuracy, which constantly expands the scope of Machine Learning applications. Dealing with such volume of data efficiently is however challenging. Fortunately, as measurements get denser, they may become gradually redundant. We can then greatly reduce the burden by finding a representation which exploits properties of the generating process and/or is tailored for the application at hand.This thesis revolves around an aspect of this idea: functional data. Data indeed consist of discrete measurements, but sometimes thinking of those as functional, we can exploit prior knowledge on smoothness to obtain a better yet lower dimensional representation. The focus is on nonlinear models for functional output regression (FOR), relying on an extension of reproducing kernel Hilbert spaces for vector-valued functions (vv-RKHS), which is the cornerstone of many nonlinear existing FOR methods. We propose to challenge those in two aspects: their computational complexity with respect to the number of measurements per function and their focusing solely on the square loss.To that end, we introduce the new framework of kernel projection learning (KPL) combining vv-RKHSs and representation of signals in dictionaries. The loss remains functional, however the model predicts only a finite number of representation coefficients. This approach retains the many advantages of vv-RKHSs yet greatly alleviates the computational burden incurred by the functional outputs. We derive two estimators in closed-form using the square loss, one for fully observed functions and one for discretized ones. We show that both are consistent in terms of excess risk. We demonstrate as well the possibility to use other differentiable and convex losses, to combine this framework with large scale kernel methods and to automatically select the dictionary using a structured penalty.In another contribution, we propose to solve the regression problem in vv-RKHSs of function-valued functions for the family of convoluted losses which we introduce. Those losses can either promote sparsity or robustness with a parameter controlling the degree of locality of those properties. Thanks to their structure, they are particularly amenable to dual approaches which we investigate. We then overcome the challenges posed by the functional nature of the dual variables by proposing two possible representations and we propose corresponding algorithms. (10.70675/4ef05483z196dz4df8z8794z240f95694566)
    DOI : 10.70675/4ef05483z196dz4df8z8794z240f95694566
  • Demonstration Of Performance For Low Cost Personal HSM
    • Urien Pascal
    , 2023, pp.879-880. This demonstration presents an original personal Hardware Secure Module (HSM) server, built from grid of secure elements and host system (Raspberry Pi), with internet connectivity. Each secure element is plugged in a board with a microcontroller providing I2C (Inter-Integrated Circuit) interface. The host system executes the open software IoSEv5 (Internet of Secure Elements version 5), which manages two T CP/IP daemons. First is used for downloading software in secure elements, second is a TLS front server that send/receive TLS packets to/from TLS backend servers running in secure elements. Applications hosted in secure elements implement a keystore, which stores cryptographic keys and computes signature over 256 bits elliptic curve. The demonstration shows the grid at work with 16 simultaneous TLS sessions performing signature operation. It shows that performance follows the Amdahl's law, with a speeding factor of about 50. (10.1109/CCNC51644.2023.10060586)
    DOI : 10.1109/CCNC51644.2023.10060586
  • Towards a centralized security architecture for SOME/IP automotive services
    • Khemissa Hamza
    • Urien Pascal
    , 2023, pp.977-978. Connected and autonomous vehicles (CAVs) consist of a number of networked computer components, called Electronic Control Units (ECUs). Scalable service-Oriented MiddlewarE over IP (SOME/IP) is a communication middleware standardized used to exchange various services between disjoint applications on distinct ECUs. However, it presents lack of authentication and confidentiality features. In this paper, we propose a centralized security architecture for SOME/IP automotive services. First, we present a lightweight symmetric cryptography based session key agreement scheme between each ECU and the manufacturer data center, which uses a random nonce, concatenation operator, a simple hash function and a keyedhash message authentication code (HMAC). Then, we define the security parameters between the different ECUs for the invehicle Ethernet-based communications. We propose the use of DTLS using pre-shared keys (DTLS-PSK) in order to secure the transmission of SOME/IP messages, (10.1109/CCNC51644.2023.10059950)
    DOI : 10.1109/CCNC51644.2023.10059950
  • Coalitional game-theoretical approach to coinvestment with application to edge computing
    • Patanè Rosario
    • Araldo Andrea
    • Chahed Tijani
    • Kiedanski Diego
    • Kofman Daniel
    , 2023, pp.517-522. We propose in this paper a coinvestment plan between several stakeholders of different types, namely a physical network owner, operating network nodes, e.g. a network operator or a tower company, and a set of service providers willing to use these resources to provide services as video streaming, augmented reality, autonomous driving assistance, etc. One such scenario is that of deployment of Edge Computing resources. Indeed, although the latter technology is ready, the high Capital Expenditure (CAPEX) cost of such resources is the barrier to its deployment. For this reason, a solid economical framework to guide the investment and the returns of the stakeholders is key to solve this issue. We formalize the coinvestment framework using coalitional game theory. We provide a solution to calculate how to divide the profits and costs among the stakeholders, taking into account their characteristics: traffic load, revenues, utility function. We prove that it is always possible to form the grand coalition composed of all the stakeholders, by showing that our game is convex. We derive the payoff of the stakeholders using the Shapley value concept, and elaborate on some properties of our game. We show our solution in simulation. (10.1109/CCNC51644.2023.10060093)
    DOI : 10.1109/CCNC51644.2023.10060093
  • Public-attention-based Adversarial Attack on Traffic Sign Recognition
    • Chi Lijun
    • Msahli Mounira
    • Memmi Gerard
    • Qiu Han
    , 2023, pp.740-745. Autonomous driving systems (ADS) can instantaneously and accurately recognize traffic signs by using deep neural networks (DNNs). Although adversarial attacks are well-known to easily fool DNNs by adding tiny but malicious perturbations, most attack methods require sufficient information about the victim models (white-box) to perform. In this paper, we propose a black-box attack in the recognition system of ADS, Public Attention Attacks (PAA), that can attack a black-box model by collecting the generic attention patterns of other white-box DNNs to transfer the attack. Particularly, we select multiple dual or triple attention patterns of white-box model combinations to generate the transferable adversarial perturbations for PAA attacks. We perform the experimentation on four well-trained models in different adversarial settings separately. The results indicate that when more white-box models the adversary collects to perform PAA, the higher the attack success rate (ASR) he can achieve to attack the target black-box model. (10.1109/CCNC51644.2023.10060485)
    DOI : 10.1109/CCNC51644.2023.10060485
  • Compressing Explicit Voxel Grid Representations: fast NeRFs become also small
    • Deng Chenxi Lola
    • Tartaglione Enzo
    , 2023, pp.1236-1245. (10.1109/WACV56688.2023.00129)
    DOI : 10.1109/WACV56688.2023.00129
  • Fuzzy Sets Methods in Image Processing and Understanding
    • Bloch Isabelle
    • Ralescu Anca
    , 2023. (10.1007/978-3-031-19425-2)
    DOI : 10.1007/978-3-031-19425-2
  • MPEG immersive video
    • Garus Patrick
    • Milovanović Marta
    • Jung Joël
    • Cagnazzo Marco
    , 2023, pp.327-356. MPEG immersive video (MIV) is a novel standard, enabling the compression of volumetric video content. In this chapter, we describe MIV, its tools, and its profiles. Given that MIV is a video-based solution, the texture and geometry information is coded using available 2D video codecs, which are independent of MIV. We present the performance of MIV with several state-of-the-art 2D codecs: VVC, AV1, and AVS3, highlighting that the eventual success of MIV does not depend on the market share of any particular 2D codec. However, using suitable tools for the coding of MIV texture or depth map atlases is an important requirement for efficient compression of immersive video. In this context, we present results related to screen content coding tools of VVC and show their potential for the compression of MIV atlases. (10.1016/B978-0-32-391755-1.00018-3)
    DOI : 10.1016/B978-0-32-391755-1.00018-3
  • Limitations of local update recovery in stabilizer-GKP codes: a quantum optimal transport approach
    • Koenig Robert
    • Rouzé Cambyse
    , 2023. Local update recovery seeks to maintain quantum information by applying local correction maps alternating with and compensating for the action of noise. Motivated by recent constructions based on quantum LDPC codes in the finite-dimensional setting, we establish an analytic upper bound on the fault-tolerance threshold for concatenated GKP-stabilizer codes with local update recovery. Our bound applies to noise channels that are tensor products of one-mode beamsplitters with arbitrary environment states, capturing, in particular, photon loss occurring independently in each mode. It shows that for loss rates above a threshold given explicitly as a function of the locality of the recovery maps, encoded information is lost at an exponential rate. This extends an early result by Razborov from discrete to continuous variable (CV) quantum systems. To prove our result, we study a metric on bosonic states akin to the Wasserstein distance between two CV density functions, which we call the bosonic Wasserstein distance. It can be thought of as a CV extension of a quantum Wasserstein distance of order 1 recently introduced by De Palma et al. in the context of qudit systems, in the sense that it captures the notion of locality in a CV setting. We establish several basic properties, including a relation to the trace distance and diameter bounds for states with finite average photon number. We then study its contraction properties under quantum channels, including tensorization, locality and strict contraction under beamsplitter-type noise channels. Due to the simplicity of its formulation, and the established wide applicability of its finite-dimensional counterpart, we believe that the bosonic Wasserstein distance will become a versatile tool in the study of CV quantum systems. (10.48550/arXiv.2309.16241)
    DOI : 10.48550/arXiv.2309.16241
  • Interband cascade technology for energy-efficient mid-infrared free-space communication
    • Didier Pierre
    • Knötig Hedwig
    • Spitz Olivier
    • Cerutti Laurent
    • Lardschneider Anna
    • Awwad Elie
    • Diaz-Thomas Daniel
    • Baranov A.
    • Weih Robert
    • Koeth Johannes
    • Schwarz Benedikt
    • Grillot Frédéric
    Photonics research, Optical Society of America, 2023, 11 (4), pp.582. Space-to-ground high-speed transmission is of utmost importance for the development of a worldwide broadband network. Mid-infrared wavelengths offer numerous advantages for building such a system, spanning from low atmospheric attenuation to eye-safe operation and resistance to inclement weather conditions. We demonstrate a full interband cascade system for high-speed transmission around a wavelength of 4.18 µm. The low-power consumption of both the laser and the detector in combination with a large modulation bandwidth and sufficient output power makes this technology ideal for a free-space optical communication application. Our proof-of-concept experiment employs a radio-frequency optimized Fabry–Perot interband cascade laser and an interband cascade infrared photodetector based on a type-II InAs/GaSb superlattice. The bandwidth of the system is evaluated to be around 1.5 GHz. It allows us to achieve data rates of 12 Gbit/s with an on–off keying scheme and 14 Gbit/s with a 4-level pulse amplitude modulation scheme. The quality of the transmission is enhanced by conventional pre- and post-processing in order to be compatible with standard error-code correction. (10.1364/PRJ.478776)
    DOI : 10.1364/PRJ.478776
  • Integration of heterogeneous components for co-simulation
    • Jerray Jawher
    • Ameur-Boulifa Rabéa
    • Apvrille Ludovic
    , 2023. Because of their complexity, embedded systems are designed with sub-systems or components taken in charge by different development teams or entities and with different modeling frameworks and simulation tools, depending on the characteristics of each component. Unfortunately, this diversity of tools and semantics makes the integration of these heterogeneous components difficult. Thus, to evaluate their integration before their hardware or software is available, one solution would be to merge them into a common modeling framework. Yet, such a holistic environment supporting many computation and computation semantics seems hard to settle. Another solution we investigate in this paper is to generically link their respective simulation environments in order to keep the strength and semantics of each component environment. The paper presents a method to simulate heterogeneous components of embedded systems in real-time. These components can be described at any abstraction level. Our main contribution is a generic glue that can analyze in real-time the state of different simulation environments and accordingly enforce the correct communication semantics between components. Once presented in a generic way, our glue is illustrated with Apache Kafka as the communication facility between simulation engines. It is then applied to two model and simulation frameworks: TTool and SystemC. Finally, Zigbee serves as a case study to illustrate the strengths of our approach.
  • Efficient learning of the structure and parameters of local Pauli noise channels
    • Rouzé Cambyse
    • Stilck Franca Daniel
    , 2023. The unavoidable presence of noise is a crucial roadblock for the development of large-scale quantum computers and the ability to characterize quantum noise reliably and efficiently with high precision is essential to scale quantum technologies further. Although estimating an arbitrary quantum channel requires exponential resources, it is expected that physically relevant noise has some underlying local structure, for instance that errors across different qubits have a conditional independence structure. Previous works showed how it is possible to estimate Pauli noise channels with an efficient number of samples in a way that is robust to state preparation and measurement errors, albeit departing from a known conditional independence structure. We present a novel approach for learning Pauli noise channels over n qubits that addresses this shortcoming. Unlike previous works that focused on learning coefficients with a known conditional independence structure, our method learns both the coefficients and the underlying structure. We achieve our results by leveraging a groundbreaking result by Bresler for efficiently learning Gibbs measures and obtain an optimal sample complexity of O(log(n)) to learn the unknown structure of the noise acting on n qubits. This information can then be leveraged to obtain a description of the channel that is close in diamond distance from O(poly(n)) samples. Furthermore, our method is efficient both in the number of samples and postprocessing without giving up on other desirable features such as SPAM-robustness, and only requires the implementation of single qubit Cliffords. In light of this, our novel approach enables the large-scale characterization of Pauli noise in quantum devices under minimal experimental requirements and assumptions. (10.48550/arXiv.2307.02959)
    DOI : 10.48550/arXiv.2307.02959
  • (Adversarial) Electromagnetic Disturbance in the Industry
    • Beckers Arthur
    • Guilley Sylvain
    • Maurine Philippe
    • O'Flynn Colin
    • Picek Stjepan
    IEEE Transactions on Computers, Institute of Electrical and Electronics Engineers, 2023, 72 (2), pp.414-422. Faults occur naturally and are responsible for reliability concerns. Faults are also an interesting tool for attackers to extract sensitive information from secure chips. In particular, non-invasive fault attacks have received a fair amount of attention. One easy way to perturb a chip without altering it is the so-called Electromagnetic Fault Injection (EMFI). Such attack has been studied in great depth, and nowadays, it is part and parcel of the state-of-the-art. Indeed, new capabilities have emerged where EM experimental benches are used to cryptanalyze chips. The progress of this "field" is fast, in terms of reproducibility, accuracy, and number of use-cases. However, there is too little awareness about such advances. In this paper, we aim to expose the true harmfulness of EMFI (including reproducibility) to enable reasonable security quotations. We also analyze protections (at hardware/firmware/system levels) in light of their efficiency. We characterize the specificity of EM fault injection compared to other injection means (laser, glitch, probing). (10.1109/TC.2022.3224373)
    DOI : 10.1109/TC.2022.3224373
  • Localization in 1D non-parametric latent space models from pairwise affinities
    • Giraud Christophe
    • Issartel Yann
    • Verzelen Nicolas
    Electronic Journal of Statistics, Shaker Heights, OH : Institute of Mathematical Statistics, 2023, 17 (1), pp.1587-1662. We consider the problem of estimating latent positions in a one-dimensional torus from pairwise affinities. The observed affinity between a pair of items is modeled as a noisy observation of a function f (x*i , x*j) of the latent positions x*i, x*j of the two items on the torus. The affinity func-tion f is unknown, and it is only assumed to fulfill some shape constraints ensuring that f(x, y) is large when the distance between x and y is small, and vice-versa. This non-parametric modeling offers a good flexibility to fit data. We introduce an estimation procedure that provably localizes all the latent positions with a maximum error of the order of log(n)/n, with high-probability. This rate is proven to be minimax optimal. A computa-tionally efficient variant of the procedure is also analyzed under some more restrictive assumptions. Our general results can be instantiated to the prob-lem of statistical seriation, leading to new bounds for the maximum error in the ordering. (10.1214/23-ejs2134)
    DOI : 10.1214/23-ejs2134
  • Nonatomic Non-Cooperative Neighbourhood Balancing Games
    • Auger David
    • Cohen Johanne
    • Lobstein Antoine
    , 2023. We introduce a game where players selfishly choose a resource and endure a cost depending on the number of players choosing nearby resources. We model the influences among resources by a weighted graph, directed or not. These games are generalizations of well-known games like Wardrop and congestion games. We study the conditions of equilibria existence and their efficiency if they exist. We conclude with studies of games whose influences among resources can be modelled by simple graphs. (10.48550/arXiv.2303.08507)
    DOI : 10.48550/arXiv.2303.08507
  • Next generation of Bluetooth and Wi-Fi networks
    • Lim Keun-Woo
    , 2023.
  • PROCÉDÉ D'ÉVALUATION DE L'ÉTAT RELATIF D'UN MOTEUR D'AÉRONEF
    • Pineau Edouard
    • Razakarivony Sébastien
    • Bonald Thomas
    , 2023.
  • Locally differentially private estimation of nonlinear functionals of discrete distributions
    • Butucea Cristina
    • Issartel Yann
    Advances in Neural Information Processing Systems, Morgan Kaufmann Publishers, 2023. We study the problem of estimating non-linear functionals of discrete distributions in the context of local differential privacy. The initial data $x_1,\ldots,x_n \in [K]$ are supposed i.i.d. and distributed according to an unknown discrete distribution $p = (p_1,\ldots,p_K)$. Only $\alpha$-locally differentially private (LDP) samples $z_1,...,z_n$ are publicly available, where the term 'local' means that each $z_i$ is produced using one individual attribute $x_i$. We exhibit privacy mechanisms (PM) that are interactive (i.e. they are allowed to use already published confidential data) or non-interactive. We describe the behavior of the quadratic risk for estimating the power sum functional $F_{\gamma} = \sum_{k=1}^K p_k^{\gamma}$, $\gamma >0$ as a function of $K, \, n$ and $\alpha$. In the non-interactive case, we study two plug-in type estimators of $F_{\gamma}$, for all $\gamma >0$, that are similar to the MLE analyzed by Jiao et al. (2017) in the multinomial model. However, due to the privacy constraint the rates we attain are slower and similar to those obtained in the Gaussian model by Collier et al. (2020). In the interactive case, we introduce for all $\gamma >1$ a two-step procedure which attains the faster parametric rate $(n \alpha^2)^{-1/2}$ when $\gamma \geq 2$. We give lower bounds results over all $\alpha$-LDP mechanisms and all estimators using the private samples.
  • Scaling by subsampling for big data, with applications to statistical learning
    • Bertail Patrice
    • Bouchouia Mohammed
    • Jelassi Ons
    • Tressou Jessica
    • Zetlaoui Mélanie
    Journal of Nonparametric Statistics, American Statistical Association, 2023, 36 (1), pp.78-117. Handling large datasets and calculating complex statistics on huge datasets require important computing resources. Using subsampling methods to calculate statistics of interest on small samples is often used in practice to reduce computational complexity, for instance using the divide and conquer strategy. In this article, we recall some results on subsampling distributions and derive a precise rate of convergence for these quantities and the corresponding quantiles. We also develop some standardization techniques based on subsampling unstandardized statistics in the framework of large datasets. It is argued that using several subsampling distributions with different subsampling sizes brings a lot of information on the behavior of statistical learning procedures: subsampling allows to estimate the rate of convergence of different algorithms, to estimate the variability of complex statistics, to estimate confidence intervals for out-of-sample errors and interpolate their values at larger scales. These results are illustrated on simulations, but also on two important datasets, frequently analyzed in the statistical learning community, EMNIST (recognition of digits) and VeReMi (analysis of Network Vehicular Reference Misbehavior). (10.1080/10485252.2023.2219782)
    DOI : 10.1080/10485252.2023.2219782