Digital data is a gold mine for modern journalism. However, datasets which interest journalists are extremely heterogeneous, ranging from highly structured (relational databases), semi-structured (JSON, XML, HTML), graphs (e.g., RDF), and text. Journalists (and other classes of users lacking advanced IT expertise, such as most non-governmental-organizations, or small public administrations) need to be able to make sense of such heterogeneous corpora, even if they lack the ability to define and deploy custom extract-transform-load workflows, especially for dynamically varying sets of data sources.
Eva Luisa Vogt, Jonathan Aristya Setyadji, Andreas Mortensen, Léa Deillon, Alejandra Inés Slagter, David Hernandez Escobar
Katie Sabrina Catherine Rosie Marsden