Sorry, you need to enable JavaScript to visit this website.
Partager

Publications

 

Les publications de nos enseignants-chercheurs sont sur la plateforme HAL :

 

Les publications des thèses des docteurs du LTCI sont sur la plateforme HAL :

 

Retrouver les publications figurant dans l'archive ouverte HAL par année :

2023

  • Integration of heterogeneous components for co-simulation
    • Jerray Jawher
    • Ameur-Boulifa Rabéa
    • Apvrille Ludovic
    , 2023. Because of their complexity, embedded systems are designed with sub-systems or components taken in charge by different development teams or entities and with different modeling frameworks and simulation tools, depending on the characteristics of each component. Unfortunately, this diversity of tools and semantics makes the integration of these heterogeneous components difficult. Thus, to evaluate their integration before their hardware or software is available, one solution would be to merge them into a common modeling framework. Yet, such a holistic environment supporting many computation and computation semantics seems hard to settle. Another solution we investigate in this paper is to generically link their respective simulation environments in order to keep the strength and semantics of each component environment. The paper presents a method to simulate heterogeneous components of embedded systems in real-time. These components can be described at any abstraction level. Our main contribution is a generic glue that can analyze in real-time the state of different simulation environments and accordingly enforce the correct communication semantics between components. Once presented in a generic way, our glue is illustrated with Apache Kafka as the communication facility between simulation engines. It is then applied to two model and simulation frameworks: TTool and SystemC. Finally, Zigbee serves as a case study to illustrate the strengths of our approach.
  • Efficient learning of the structure and parameters of local Pauli noise channels
    • Rouzé Cambyse
    • Stilck Franca Daniel
    , 2023. The unavoidable presence of noise is a crucial roadblock for the development of large-scale quantum computers and the ability to characterize quantum noise reliably and efficiently with high precision is essential to scale quantum technologies further. Although estimating an arbitrary quantum channel requires exponential resources, it is expected that physically relevant noise has some underlying local structure, for instance that errors across different qubits have a conditional independence structure. Previous works showed how it is possible to estimate Pauli noise channels with an efficient number of samples in a way that is robust to state preparation and measurement errors, albeit departing from a known conditional independence structure. We present a novel approach for learning Pauli noise channels over n qubits that addresses this shortcoming. Unlike previous works that focused on learning coefficients with a known conditional independence structure, our method learns both the coefficients and the underlying structure. We achieve our results by leveraging a groundbreaking result by Bresler for efficiently learning Gibbs measures and obtain an optimal sample complexity of O(log(n)) to learn the unknown structure of the noise acting on n qubits. This information can then be leveraged to obtain a description of the channel that is close in diamond distance from O(poly(n)) samples. Furthermore, our method is efficient both in the number of samples and postprocessing without giving up on other desirable features such as SPAM-robustness, and only requires the implementation of single qubit Cliffords. In light of this, our novel approach enables the large-scale characterization of Pauli noise in quantum devices under minimal experimental requirements and assumptions. (10.48550/arXiv.2307.02959)
    DOI : 10.48550/arXiv.2307.02959
  • (Adversarial) Electromagnetic Disturbance in the Industry
    • Beckers Arthur
    • Guilley Sylvain
    • Maurine Philippe
    • O'Flynn Colin
    • Picek Stjepan
    IEEE Transactions on Computers, Institute of Electrical and Electronics Engineers, 2023, 72 (2), pp.414-422. Faults occur naturally and are responsible for reliability concerns. Faults are also an interesting tool for attackers to extract sensitive information from secure chips. In particular, non-invasive fault attacks have received a fair amount of attention. One easy way to perturb a chip without altering it is the so-called Electromagnetic Fault Injection (EMFI). Such attack has been studied in great depth, and nowadays, it is part and parcel of the state-of-the-art. Indeed, new capabilities have emerged where EM experimental benches are used to cryptanalyze chips. The progress of this "field" is fast, in terms of reproducibility, accuracy, and number of use-cases. However, there is too little awareness about such advances. In this paper, we aim to expose the true harmfulness of EMFI (including reproducibility) to enable reasonable security quotations. We also analyze protections (at hardware/firmware/system levels) in light of their efficiency. We characterize the specificity of EM fault injection compared to other injection means (laser, glitch, probing). (10.1109/TC.2022.3224373)
    DOI : 10.1109/TC.2022.3224373
  • Localization in 1D non-parametric latent space models from pairwise affinities
    • Giraud Christophe
    • Issartel Yann
    • Verzelen Nicolas
    Electronic Journal of Statistics, Shaker Heights, OH : Institute of Mathematical Statistics, 2023, 17 (1), pp.1587-1662. We consider the problem of estimating latent positions in a one-dimensional torus from pairwise affinities. The observed affinity between a pair of items is modeled as a noisy observation of a function f (x*i , x*j) of the latent positions x*i, x*j of the two items on the torus. The affinity func-tion f is unknown, and it is only assumed to fulfill some shape constraints ensuring that f(x, y) is large when the distance between x and y is small, and vice-versa. This non-parametric modeling offers a good flexibility to fit data. We introduce an estimation procedure that provably localizes all the latent positions with a maximum error of the order of log(n)/n, with high-probability. This rate is proven to be minimax optimal. A computa-tionally efficient variant of the procedure is also analyzed under some more restrictive assumptions. Our general results can be instantiated to the prob-lem of statistical seriation, leading to new bounds for the maximum error in the ordering. (10.1214/23-ejs2134)
    DOI : 10.1214/23-ejs2134
  • Nonatomic Non-Cooperative Neighbourhood Balancing Games
    • Auger David
    • Cohen Johanne
    • Lobstein Antoine
    , 2023. We introduce a game where players selfishly choose a resource and endure a cost depending on the number of players choosing nearby resources. We model the influences among resources by a weighted graph, directed or not. These games are generalizations of well-known games like Wardrop and congestion games. We study the conditions of equilibria existence and their efficiency if they exist. We conclude with studies of games whose influences among resources can be modelled by simple graphs. (10.48550/arXiv.2303.08507)
    DOI : 10.48550/arXiv.2303.08507
  • Next generation of Bluetooth and Wi-Fi networks
    • Lim Keun-Woo
    , 2023.
  • PROCÉDÉ D'ÉVALUATION DE L'ÉTAT RELATIF D'UN MOTEUR D'AÉRONEF
    • Pineau Edouard
    • Razakarivony Sébastien
    • Bonald Thomas
    , 2023.
  • Solving stochastic weak Minty variational inequalities without increasing batch size
    • Pethick Thomas
    • Fercoq Olivier
    • Latafat Puya
    • Patrinos Panagiotis
    • Cevher Volkan
    , 2023. This paper introduces a family of stochastic extragradient-type algorithms for a class of nonconvex-nonconcave problems characterized by the weak Minty variational inequality (MVI). Unlike existing results on extragradient methods in the monotone setting, employing diminishing stepsizes is no longer possible in the weak MVI setting. This has led to approaches such as increasing batch sizes per iteration which can however be prohibitively expensive. In contrast, our proposed methods involves two stepsizes and only requires one additional oracle evaluation per iteration. We show that it is possible to keep one fixed stepsize while it is only the second stepsize that is taken to be diminishing, making it interesting even in the monotone setting. Almost sure convergence is established and we provide a unified analysis for this family of schemes which contains a nonlinear generalization of the celebrated primal dual hybrid gradient algorithm.
  • A Consistent Diffusion-Based Algorithm for Semi-Supervised Graph Learning
    • Bonald Thomas
    • de Lara Nathan
    , 2023. The task of semi-supervised classification aims at assigning labels to all nodes of a graph based on the labels known for a few nodes, called the seeds. One of the most popular algorithms relies on the principle of heat diffusion, where the labels of the seeds are spread by thermoconductance and the temperature of each node at equilibrium is used as a score function for each label. In this paper, we prove that this algorithm is not consistent unless the temperatures of the nodes at equilibrium are centered before scoring. This crucial step does not only make the algorithm provably consistent on a block model but brings significant performance gains on real graphs.
  • ADT: AI-Driven network Telemetry processing on routers
    • Foroughi Parisa
    • Brockners Frank
    • Rougier Jean-Louis
    Computer Networks, Elsevier, 2023, 220, pp.109474-1:109474-17. Network monitoring is a pivotal part of network management and operations. It is responsible for monitoring the behavior of the network to assure its functionality within expectation and to guarantee a smooth-running environment for enabling of various services. Therefore, operators are interested in gaining a comprehensive assessment of their network elements and tracking operational changes to facilitate timely correction of any deviation. Commonly, this assessment is achieved by performing regular manual checks of different operational counters and defining expert rules from known root causes. The common approach requires the maintenance of a regularly updated set of rules and only goes as far as the operator's pre-gained knowledge of the system. With the growing complexity of the networks as well as the availability of more data, a more efficient monitoring approach is necessary to address the emerging network monitoring requirements. In this paper, a novel unsupervised approach is proposed that is capable of exploring a broader set of counters (not limited to the handpicked Key Performance Indicators (KPIs)). The goal is to leverage the dependencies between the counters in order to discover complex state changes that might have otherwise slipped the operator's view. This paper proposes ADT, an AI-driven telemetry processing solution that facilitates monitoring of a larger set of counters. The Detector block of ADT is known as DESTIN, a multivariate unsupervised change detection for high dimensional time-series data of originally low effective dimension, which provides near real-time state assessment of network devices. The efficiency of the proposed approach is demonstrated and compared with well-known methodologies on an experimental test-bed. The method's performance is also explored extensively considering different criteria such as traffic type, device and the type of events to identify its potentials and limitations. The datasets used for the evaluation are made publicly available. (10.1016/j.comnet.2022.109474)
    DOI : 10.1016/j.comnet.2022.109474
  • DNA code from cyclic and skew cyclic codes over F 4 [v]/⟨v 3 ⟩
    • Prakash Om
    • Singh Ashutosh
    • Verma Ram Krishna
    • Solé Patrick
    • Cheng Wei
    Entropy, MDPI, 2023. The main motivation of this work is to study and obtain some reversible and DNA codes of length n with better parameters. Here, we first investigate the structure of cyclic and skew cyclic codes over the chain ring R : = F 4 [ v ] / ⟨ v 3 ⟩ . We show an association between the codons and the elements of R using a Gray map. Under this Gray map, we study reversible and DNA codes of length n. Finally, several new DNA codes are obtained that have improved parameters than previously known codes. We also determine the Hamming and the Edit distances of these codes. (10.3390/e25020239)
    DOI : 10.3390/e25020239
  • Quadratic error bound of the smoothed gap and the restarted averaged primal-dual hybrid gradient
    • Fercoq Olivier
    Open Journal of Mathematical Optimization, Centre Mersenne, 2023, 4, pp.1-34. We study the linear convergence of the primal-dual hybrid gradient method. After a review of current analyses, we show that they do not explain properly the behavior of the algorithm, even on the most simple problems. We thus introduce the quadratic error bound of the smoothed gap, a new regularity assumption that holds for a wide class of optimization problems. Equipped with this tool, we manage to prove tighter convergence rates. Then, we show that averaging and restarting the primal-dual hybrid gradient allows us to leverage better the regularity constant. Numerical experiments on linear and quadratic programs, ridge regression and image denoising illustrate the findings of the paper. (10.5802/ojmo.26)
    DOI : 10.5802/ojmo.26
  • Proceedings of The Semantic Web: ESWC 2023 Satellite Events
    • Pesquita Catia
    • Skaf-Molli Hala
    • Efthymiou Vasilis
    • Kirrane Sabrina
    • Ngonga Axel
    • Collarana Diego
    • Cerqueira Renato
    • Alam Mehwish
    • Trojahn Cassia
    • Hertling Sven
    , 2023, 13998. (10.1007/978-3-031-43458-7)
    DOI : 10.1007/978-3-031-43458-7
  • Affine invariant integrated rank-weighted statistical depth: properties and finite sample analysis
    • Clémençon Stephan
    • Mozharovskyi Pavlo
    • Staerman Guillaume
    Electronic Journal of Statistics, Shaker Heights, OH : Institute of Mathematical Statistics, 2023, 17 (2), pp.3854 - 3892. Because it determines a center-outward ordering of observations in Rd with d≥2, the concept of statistical depth permits to define quantiles and ranks for multivariate data and use them for various statistical tasks (e.g. inference, hypothesis testing). Whereas many depth functions have been proposed ad-hoc in the literature since the seminal contribution of [50], not all of them possess the properties desirable to emulate the notion of quantile function for univariate probability distributions. In this paper, we propose an extension of the integrated rank-weighted statistical depth (IRW depth in abbreviated form) originally introduced in [40], modified in order to satisfy the property of affine invariance, fulfilling thus all the four key axioms listed in the nomenclature elaborated by [59]. The variant we propose, referred to as the affine invariant IRW depth (AI-IRW in short), involves the precision matrix of the (supposedly square integrable) d-dimensional random vector X under study, in order to take into account the directions along which X is most variable to assign a depth value to any point x∈Rd. The accuracy of the sampling version of the AI-IRW depth is investigated from a non-asymptotic perspective. Namely, a concentration result for the statistical counterpart of the AI-IRW depth is proved. Beyond the theoretical analysis carried out, applications to anomaly detection are considered and numerical results are displayed, providing strong empirical evidence of the relevance of the depth function we propose here. (10.1214/23-EJS2189)
    DOI : 10.1214/23-EJS2189
  • Dynamic Autoencoders Against Adversarial Attacks
    • Chabanne Hervé
    • Despiegel Vincent
    • Gentric Stéphane
    • Guiga Linda
    Procedia Computer Science, Elsevier, 2023, 220, pp.782-787. Neural Networks are the target of numerous adversarial attacks. In those, the adversary perturbs a model's input with a noise that is small, but large enough to fool the model. In this article, we propose to dynamically add autoencoders from a pretrained set to a base model as a countermeasure to such attacks. This doing, we modify the underlying labels regions of the model to be protected, letting the adversary unable to craft relevant adversarial perturbations. Our experiments confirm the efficiency of our protection when the pretrained set has enough elements. (10.1016/j.procs.2023.03.104)
    DOI : 10.1016/j.procs.2023.03.104
  • Improving Causal Learning Scalability and Performance using Aggregates and Interventions
    • Fadiga Kanvaly
    • Houzé Etienne
    • Diaconescu Ada
    • Dessalles Jean-Louis
    ACM Transactions on Autonomous and Adaptive Systems, Association for Computing Machinery (ACM), 2023. Smart homes are Cyber-Physical Systems (CPS) where multiple devices and controllers cooperate to achieve high-level goals. Causal knowledge on relations between system entities is essential for enabling system self-adaption to dynamic changes. As house configurations are diverse, this knowledge is difficult to obtain. In previous work, we proposed to generate Causal Bayesian Networks (CBN) as follows. Starting with considering all possible relations, we progressively discarded non-correlated variables. Next, we identified causal relations from the remaining correlations by employing “ do-operations ”. The obtained CBN could then be employed for causal inference. The main challenges of this approach included: “non-doable variables” and limited scalability. To address these issues, we propose three extensions: i) early pruning weakly correlated relations to reduce the number of required do-operations; ii) introducing aggregate variables that summarize relations between weakly-coupled sub-systems; iii) applying the method a second time to perform indirect do interventions and handle non-doable relations. We illustrate and evaluate the efficiency of these contributions via examples from the smart home and power grid domain. Our proposal leads to a decrease in the number of operations required to learn the CBN and in an increased accuracy of the learned CBN, paving the way towards applications in large CPS. (10.1145/3607872)
    DOI : 10.1145/3607872
  • Some Complexity Considerations on the Uniqueness of Graph Colouring
    • Hudry Olivier
    • Lobstein Antoine
    WSEAS Transactions on Mathematics, World Scientific and Engineering Academy and Society (WSEAS), 2023, 22, pp.art. #54, 483-493. For some well-known NP-complete problems, linked to the satisfiability of Boolean formulas and the colourability of a graph, we study the variation which consists in asking about the uniqueness of a solution.In particular, we show that five decision problems, Unique Satisfiability (U-SAT), Unique k-Satisfiability (U-k-SAT), k ≥ 3, Unique One-in-Three Satisfiability (U-1-3-SAT), Unique k-Colouring (U-k-COL), k ≥ 3, and Unique Colouring (U-COL), have equivalent complexities, up to polynomials —when dealing with colourings, we forbid permutations ofcolours. As a consequence, all are NP-hard and belong to the class DP. We also consider the problems U-2-SAT, U-2-COL and Unique Optimal Colouring (U-OCOL). (10.37394/23206.2023.22.54)
    DOI : 10.37394/23206.2023.22.54
  • Towards a Development Process for Multi-CPU Distributed Synchronous Software Applications
    • Lubat Eric
    • Jenn Eric
    • Blouin Dominique
    • Kaufmann Marc
    Psychology in Spain, Colegio Oficial de Psicólogos, 2023, pp.549-558. (10.1109/MODELS-C59198.2023.00092)
    DOI : 10.1109/MODELS-C59198.2023.00092
  • Online Matching in Geometric Random Graphs
    • Sentenac Flore
    • Noiry Nathan
    • Lerasle Matthieu
    • Ménard Laurent
    • Perchet Vianney
    , 2023. We investigate online maximum cardinality matching, a central problem in ad allocation. In this problem, users are revealed sequentially, and each new user can be paired with any previously unmatched campaign that it is compatible with. Despite the limited theoretical guarantees, the greedy algorithm, which matches incoming users with any available campaign, exhibits outstanding performance in practice. Some theoretical support for this practical success has been established in specific classes of graphs, where the connections between different vertices lack strong correlations-an assumption not always valid in real-world situations. To bridge this gap, we focus on the following model: both users and campaigns are represented as points uniformly distributed in the interval [0, 1], and a user is eligible to be paired with a campaign if they are "similar enough," meaning the distance between their respective points is less than c/N , where c > 0 is a model parameter. As a benchmark, we determine the size of the optimal offline matching in these bipartite random geometric graphs. We achieve this by introducing an algorithm that constructs the optimal matching and analyzing it. We then turn to the online setting and investigate the number of matches made by the online algorithm CLOSEST, which pairs incoming points with their nearest available neighbors in a greedy manner. We demonstrate that the algorithm's performance can be compared to its fluid limit, which is completely characterized as the solution to a specific partial differential equation (PDE). From this PDE solution, we can compute the competitive ratio of CLOSEST, and our computations reveal that it remains significantly better than its worst-case guarantee. This model turns out to be closely related to the online minimum cost matching problem, and we can extend the results obtained here to refine certain findings in that area of research. Specifically, we determine the exact asymptotic cost of CLOSEST in the ϵ-excess regime, providing a more accurate estimate than the previously known loose upper bound.
  • Manifold Learning via Linear Tangent Space Alignment (LTSA) for Accelerated Dynamic MRI With Sparse Sampling
    • Djebra Yanis
    • Marin Thibault
    • Han Paul K
    • Bloch Isabelle
    • Fakhri Georges El
    • Ma Chao
    IEEE Transactions on Medical Imaging, Institute of Electrical and Electronics Engineers, 2023, 42 (1), pp.158-169. The spatial resolution and temporal framerate of dynamic magnetic resonance imaging (MRI) can be improved by reconstructing images from sparsely sampled k-space data with mathematical modeling of the underlying spatiotemporal signals. These models include sparsity models, linear subspace models, and non-linear manifold models. This work presents a novel linear tangent space alignment (LTSA) model-based framework that exploits the intrinsic low-dimensional manifold structure of dynamic images for accelerated dynamic MRI. The performance of the proposed method was evaluated and compared to stateof-the-art methods using numerical simulation studies as well as 2D and 3D in vivo cardiac imaging experiments. The proposed method achieved the best performance in image reconstruction among all the compared methods. The proposed method could prove useful for accelerating many MRI applications, including dynamic MRI, multi-parametric MRI, and MR spectroscopic imaging. (10.1109/TMI.2022.3207774)
    DOI : 10.1109/TMI.2022.3207774
  • Test your samples jointly: Pseudo-reference for image quality evaluation
    • Tworski Marcelin
    • Lathuilière Stéphane
    , 2023. In this paper, we address the well-known image quality assessment problem but in contrast from existing approaches that predict image quality independently for every images, we propose to jointly model different images depicting the same content to improve the precision of quality estimation. This proposal is motivated by the idea that multiple distorted images can provide information to disambiguate image features related to content and quality. To this aim, we combine the feature representations from the different images to estimate a pseudo-reference that we use to enhance score prediction. Our experiments show that at test-time, our method successfully combines the features from multiple images depicting the same new content, improving estimation quality.
  • Zero-shot spatial layout conditioning for text-to-image diffusion models
    • Couairon Guillaume
    • Careil Marlène
    • Cord Matthieu
    • Lathuilière Stéphane
    • Verbeek Jakob
    , 2023. Large-scale text-to-image diffusion models have significantly improved the state of the art in generative image modelling and allow for an intuitive and powerful user interface to drive the image generation process. Expressing spatial constraints, e.g. to position specific objects in particular locations, is cumbersome using text; and current text-based image generation models are not able to accurately follow such instructions. In this paper we consider image generation from text associated with segments on the image canvas, which combines an intuitive natural language interface with precise spatial control over the generated content. We propose ZestGuide, a zero-shot segmentation guidance approach that can be plugged into pre-trained text-to-image diffusion models, and does not require any additional training. It leverages implicit segmentation maps that can be extracted from cross-attention layers, and uses them to align the generation with input masks. Our experimental results combine high image quality with accurate alignment of generated content with input segmentations, and improve over prior work both quantitatively and qualitatively, including methods that require training on images with corresponding segmentations. Compared to Paint with Words, the previous state-of-the art in image generation with zero-shot segmentation conditioning, we improve by 5 to 10 mIoU points on the COCO dataset with similar FID scores.
  • Visualization Empowerment: How to Teach and Learn Data Visualization
    • Bach Benjamin
    • Carpendale Sheelagh
    • Hinrichs Uta
    • Huron Samuel
    , 2023, pp.10.4230/DagRep.12.6.83. Data visualization is becoming an important asset for a data-literate, informed, and critical society. Despite the variety of existing resources to teach theories and practical skills in this domain, little is known about 1) how learning processes in the context of visualization unfold and 2) best practices for engaging and teaching data visualization to diverse audiences and in different contexts. This Dagstuhl Seminar invited practitioners, researchers, and teachers from the areas of visualization, design, education and cognitive psychology to explore these questions from multiple perspectives. Through a range of practical activities, talks, and discussions, we have begun characterizing and classifying teaching methodologies. We have redacted a pedagogical manifesto, and started formalizing the concept of improvisation with visualization in the context of teaching and learning. We have also interrogated creativity as an important aspect of visualization teaching and learning and explored links between data physicalization and visualization teaching activities. Across these different themes, we have begun to map out the challenges of visualization teaching and learning and the opportunities for research and practice in this area. (10.4230/DagRep.12.6.83)
    DOI : 10.4230/DagRep.12.6.83
  • Interactive Depixelization of Pixel Art through Spring Simulation
    • Matusovic Marko
    • Parakkat Amal Dev
    • Eisemann Elmar
    Computer Graphics Forum, Wiley, 2023, 42 (2). We introduce an approach for converting pixel art into high-quality vector images. While much progress has been made on automatic conversion, there is an inherent ambiguity in pixel art, which can lead to a mismatch with the artist’s original intent. Further, there is room for incorporating aesthetic preferences during the conversion. In consequence, this work introduces an interactive framework to enable users to guide the conversion process towards high-quality vector illustrations. A key idea of the method is to cast the conversion process into a spring-system optimization that can be influenced by the user. Hereby, it is possible to resolve various ambiguities that cannot be handled by an automatic algorithm.
  • The Software Heritage License Dataset (2022 Edition)
    • González-Barahona Jesús M.
    • Montes-Leon Sergio
    • Robles Gregorio
    • Zacchiroli Stefano
    Empirical Software Engineering, Springer Verlag, 2023. Context: When software is released publicly, it is common to include with it either the full text of the license or licenses under which it is published, or a detailed reference to them. Therefore public licenses, including FOSS (free, open source software) licenses, are usually publicly available in source code repositories. Objective: To compile a dataset containing as many documents as possible that contain the text of software licenses, or references to the license terms. Once compiled, characterize the dataset so that it can be used for further research, or practical purposes related to license analysis. Method: Retrieve from Software Heritage-the largest publicly available archive of FOSS source code-all versions of all files whose names are commonly used to convey licensing terms. All retrieved documents will be characterized in various ways, using automated and manual analyses. Results: The dataset consists of 6.9 million unique license files. Additional metadata about shipped license files is also provided, making the dataset ready to use in various contexts, including: file length measures, MIME type, SPDX license (detected using ScanCode), and oldest appearance. The results of a manual analysis of 8102 documents is also included, providing a ground truth for further analysis. The dataset is released as open data as an archive file containing all deduplicated license files, plus several portable CSV files with metadata, referencing files via cryptographic checksums. Conclusions: Thanks to the extensive coverage of Software Heritage, the dataset presented in this paper covers a very large fraction of all software licenses for public code. We have assembled a large body of software licenses, characterized it quantitatively and qualitatively, and validated that it is mostly composed of licensing information and includes almost all known license texts. The dataset can be used to conduct empirical studies on open source licensing, training of automated license classifiers, natural language processing (NLP) analyses of legal texts, as well as historical and phylogenetic studies on FOSS licensing. It can also be used in practice to improve tools detecting licenses in source code. (10.1007/s10664-023-10377-w)
    DOI : 10.1007/s10664-023-10377-w