The debsources dataset: Two decades of debian source code metadata

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

We present the Debsources Dataset: distribution metadata and source code metrics spanning two decades of Free and Open Source Software (FOSS) history, seen through the lens of the Debian distribution. Debsources is a software platform used to gather, search, and publish on the Web the full source code of the Debian operating system, as well as measures about it. A notable public instance of Debsources is available at http://sources.debian.net, it includes both current and historical releases of Debian. Plugins to compute popular source code metrics (lines of code, defined symbols, disk usage) and other derived data (e.g., Checksums) have been written, integrated, and run on all the source code available on sources.debian.net. The Debsources Dataset is a PostgreSQL database dump of sources.debian.net metadata, as of February 10th, 2015. The dataset contains both Debian-specific metadata - e.g., which software packages are available in which release, which source code file belong to which package, release dates, etc. - and source code information gathered by running Debsources plugins. The Debsources Dataset offer a very long-term historical view of the macro-level evolution and constitution of FOSS through the lens of popular, representative FOSS projects of their times.

Original languageEnglish
Title of host publicationProceedings - 12th Working Conference on Mining Software Repositories, MSR 2015
PublisherIEEE Computer Society
Pages466-469
Number of pages4
ISBN (Electronic)9780769555942
DOIs
Publication statusPublished - 4 Aug 2015
Externally publishedYes
Event12th Working Conference on Mining Software Repositories, MSR 2015, co-located with the 37th ACM/IEEE International Conference on Software Engineering, ICSE 2015 - Florence, Italy
Duration: 16 May 201517 May 2015

Publication series

NameIEEE International Working Conference on Mining Software Repositories
Volume2015-August
ISSN (Print)2160-1852
ISSN (Electronic)2160-1860

Conference

Conference12th Working Conference on Mining Software Repositories, MSR 2015, co-located with the 37th ACM/IEEE International Conference on Software Engineering, ICSE 2015
Country/TerritoryItaly
CityFlorence
Period16/05/1517/05/15

Keywords

  • Debian
  • Free software
  • Open source
  • Software evolution
  • Source code

Fingerprint

Dive into the research topics of 'The debsources dataset: Two decades of debian source code metadata'. Together they form a unique fingerprint.

Cite this