Data Preservation At CERN

CERN data preseervation

View of CERN in Geneva, Switzerland.

My European travels have continued after my visit to EGU, last week I was in Geneva at CERN, also known as the European Organization for Nuclear Research. I was attending the PV2023 conference titled ‘Adding value (to) and preserving Scientific & Technical data’. This conference was an event on a very different scale, with around only 100 people attending and two parallel sessions. The aim of the event was to address prospects of data preservation, stewardship and value-adding to scientific data and research related information.

The conference began on Tuesday afternoon, 2nd May, with a series of plenary talks focused on ensuring long-term data and knowledge preservation (the “P” in PV). The day concluded with a ‘Posters minute madness’ slot where all the poster presenters had a minute to give an overview of their work. It worked really well and, amazingly, people kept to time and it gave an interesting opportunity to see a range of work in a short period, and then we could go and look at the posters themselves during the networking event and breaks that followed.

I was attending the conference on behalf of Telespazio UK, presenting the work of the Instrument Data Evaluation and Analysis Service Quality Assurance for Earth Observation, known as IDEAS-QA4EO. My talk was in a parallel session on Wednesday, a session which focused on adding value to data and facilitating data use (the “V” in PV). My presentation was entitled Reprocessing and Quality Control of Heritage Third Party Optical Earth Observation Missions, about reprocessing historical optical Earth Observation (EO) missions with a focus on the Marine Observation Satellite-1 (MOS-1). The presentation can be found here. The conference ended with on the Thursday with a series of plenary talks on both preservation and valorisation that included a focus on how to make repositories trustworthy, alongside setting up platforms, preservation utilities and datasets to make their use by scientists easier.

The conference covered space science and EO, and highlighted the need to make datasets available to those already within these communities and those new entrants who may have innovative uses for the data. To enable this to happen requires careful preservation and curation of the data alongside making online search and usage more straightforward, and this can problematic as it’s unclear how innovative users will want to perform these searches. As well as considering datasets, it’s also essential to preserve workflows, i.e., the software used and the overall processing environment. There was also a thought-provoking talk on the carbon footprint of the FAIR data model (Findability, Accessibility, Interoperability, and Reusability), suggesting we need to consider the carbon cost of the data decisions we make. The posters and oral presentations, many with associated short conference papers, can be found online.

Synchrocyclotron

Synchrocyclotron at CERN.

Whilst at CERN I had to take the opportunity to tour the Synchrocyclotron, CERN’s first accelerator that was constructed as CERN itself was being built. Coming into operation in 1957 it provided beams for CERN’s first experiments in particle and nuclear physics and it was fascinating to look around. Today, of course, CERN, operates the Large Hadron Collider and has approximately 300 PetaBytes of data stored in the on-site archive. This data was also important to the conference as, excitingly, we were all given all CERN data drives populated with open-source data – at least, I was excited about getting it, even though Andy saw it as just another thing to dust when I brought it home!

I am glad to say that I’m back home now, and have a bit of time to recover before heading off to North America in June for several Open Geospatial Consortium (OGC) events.

Leave a Reply

Your email address will not be published. Required fields are marked *

Time limit is exhausted. Please reload CAPTCHA.