First objective of the JISC-supported Sonex initiative was to identify and analyse deposit opportunities (use cases) for ingest of research papers (and potentially other scholarly work) into repositories. Later on, the project scope widened to include identification and dissemination of various projects being developed at institutions in relation to the deposit usecases previously analyzed. Finally, Sonex was recently asked to extend its analysis of deposit opportunities to research data.






Wednesday, 27 April 2011

National initiatives for promoting data management strategies: an overview


- "Hello, I want to deposit my data"
- "Sir, this is a library!"
- "Sorry" -he whispers- "I want to deposit my data".
(as told by Brian Hole, British Library, along his presentation of the DRYAD UK initiative)


  Main objective of the JISC MRD International Workshop held last month was to review progress achieved by the JISC Managing Research Data Programme and to discuss this in the context of broader international developments.

As stated in the workshop programme overview, "this dimension reflects key partnerships which JISC, the JISCMRD Programme and the DCC has been building through the IDCC Conference, the Knowledge Exchange and other initiatives. They include the Australian National Data Service, the NSF funded DataNet Projects, institutions in the US and Australia, the DFG, SURF, DANS etc".

Whithin the broader context, besides a couple of preliminary talks on the European Union approach to (and future funding of) data management initiatives -by John Wood, on the EU 'Riding the wave' report, and by Carlos Morais-Pires on the Digital Agenda for Europe- the workshop featured a specific session on "National and international infrastructure initiatives" whose first panel was called "Approaches and strategies in the UK, US, and Germany". Australian and Dutch national or specific approaches were also discussed, either at this session or later along the event.

Besides the national initiatives featured in this and further sessions along the meeting -it was reassuring to see such a broad scope of strategies or already running projects taking place at the same time in so many different countries- there are also additional, sometimes preliminary initiatives for promoting data management policies at national or institutional level in other countries such as Finland, Portugal, France, Poland or South Africa.

As new initiatives for research data management keep steadily coming up, this session was an opportunity to get an informal update on DCC's report 'Comparative Study of International Approaches to Enabling the Sharing of Research Data' - see its summary and main findings here as of Nov 2008.

Digital Curation Centre - UK
Kevin Ashley, Digital Curation Centre (DCC), described the present picture of data management in the UK as "a new context", where Universities are increasingly willing to take responsibility for data management (specially in areas not covered by Data Centres).
Once UK funder and NSF rules for Data Management Planning are being implemented, this in-advance planning is becoming very important for funders, researchers, institutions, collaborators and reusers. DCC current tasks include integrating different Data Discovery Services plus building institutional capacities: skills, policies, etc. Besides that, DCC is providing the new DMP Online service aimed to produce and maintain Data Management Plans.
Good news is that, despite varying degrees of involvement, institutions in the UK have accepted their role in RDM.

NSF-funded DataNet Projects - US
A summary of present state of research data management in the US was provided by presentations of the DataONE and DataConservancy initiatives, resp. delivered by William Michener (University Libraries at U New Mexico) and Sayeed Choudhury (Johns Hopkins University).

After stating that "researchers are presently using 90% of their time managing data instead of interpreting them", W. Michener presented the Data Observation Network for Earth (DataONE) initiative (a live DataONE presentation at U of Tennessee is available). This NSF-supported initiative aims to ensure preservation and access to multi-scale, multi-discipline, and multi-national science data. DataONE Coordinating Nodes around the world will help achieving needed international collaboration for solving the grand science and data challenges, particularly with regard to education.

The DataConservancy initiative aims to research, design, implement, deploy, and sustain data curation infrastructure for cross-disciplinary discovery with an emphasis on observational data. S. Choudhury's presentation stressed the need for data preservation as a necessary condition for data reuse and introduced the recent connection of data and publications through arXiv.org as one of the pilot projects that build upon the Project APIs.

DFG - Germany
New DFG information infrastructure projects in Germany were presented by Dr Stefan Winkler-Nees, who mentioned both Jan 2009 DFG Recommendations for Secure Storage and Availability of Digital Primary Research Data, as a base report for promoting standardized work in the data management area, and DFG running call for proposals "Information infrastructures for research data". Selected projects at this call are due to be shortly announced and will start on May/Jun'2011. Finally, in a a common line of thought with other initiatives, Dr. Winkler-Nees mentioned DFG is aiming for teaching and qualification of both researchers and data curators.


SURF Foundation & DANS - The Netherlands
Later on along the workshop, John Doove presented the SURF Enhanced Publications initiative within the SURFshare programme 2007-2011. Six new projects funded along 2011 by the SURF Foundation will allow researchers from a variety of disciplines to share datasets, illustrations, audio files, and musical scores with fellow researchers in the context of Enhanced Publications (programme video available on YouTube). There were already two previous grants rounds for Enhanced Publications. The six running projects, whose results are due in May 2011, take place within five disciplines: Economics (Open Data and Publications, Tilburg University), Linguistics (Lenguas de Bolivia, Radboud University Nijmegen, and Enhanced NIAS Publications, KNAW-Royal Netherlands Academy of Arts ans Sciences), Musicology (The Other Josquin, University Utrecht), Communication sciences (Enhancing Scholarly Publishing in the Humanities and Social Sciences, KNAW) and Geosciences (VPcross, KNAW).

The Dutch strategy for increasing research data available online was completed with the presentation "Sustainable and Trusted Data Management" delivered by Laurent Sesink (DANS-Data Archiving and Networked Services). DANS, est. 2005, deals with storage and continuous accessibility of research data in
the social sciences and humanities and promotes the 'Data Seal of Approval' for certification of data repositories, guaranteeing via a series of required criteria a qualitatively high and reliable way of managing research data.

Australian National Data Service (ANDS) - Australia
Finally, Andrew Treloar, Director of Technology, Australian National Data Service (ANDS), supplied a comprehensive perspective from a national infrastructure provider and in a way summarized previous talks by saying that, despite differences, there are common themes emerging in national approaches to data management, as there are things only they can do. Along his plenary presentation "Data: Its origins in the past, what the problems are in the present, and how national responses can help fix the future" he mentioned for instance that Hubble Space Telescope-related publication statistics show double research is being done thanks to data reuse. Efficiency, validation, integrity of scholarly records, value for money and self-interest were listed as (non-altruistic) arguments for data reuse.

Having the chance to attend this series of brilliant presentations and checking out how policies for opening access to research data keep spreading over institutions and countries were undoubtedly part of the Birmingham workshop highlights. Next opportunity for keeping up with it all will be next November at the Knowledge Exchange Workshop on Research Data Management in Bonn, Germany.

Monday, 25 April 2011

Could external cooperation improve collection of specific JISC MRD project-related information?


  In forthcoming days SONEX will be publishing some posts on the JISC MRD Programme International Workshop held last March 28-29th at Aston Business School Conference Centre, Birmingham. Certain aspects debated at this comprehensive meeting were very useful for establishing an approach for dealing with research data management from a SONEX viewpoint, as debated in a SONEX meeting at EDINA on Mar 30th whose outcome will also be shortly blogged.

See IUCr Brian McMahon's report for a general review on the JISC MRD workshop.

One of the most visible disciplinary approaches to data management presented at the JISC MRD event -which featured all kinds of institutional and subject-based initiatives in the area- was the one coming from meteorology, palaeoclimatology and climate-related sciences: there was a presentation of the PEG-BOARD Project (U of Bristol) at the Subject-Oriented Approaches session on Monday, followed by ACRID (U of East Anglia & STFC) and Metafor (BADC & STFC) Project presentations on Tuesday afternoon.

One of the most relevant features of these climate-related projects is interdisciplinarity. PEG-BOARD Project in particular aims to serve the archaeology research community by supplying them their paleoclimate data.


A few specific aspects about PEG-BOARD were discussed after the project presentation. Interesting thing about them is they were not mentioned along the talk, nor are they reported at the project site:

- Due to the project interdisciplinarity, there are two clearly different user groups for palaeoclimatology data produced: climatologists, who will understand the nature of involved datasets, as they're central to their discipline, and archaeologists, who don't and need not know much about the data format but need the information contained in it for their own purposes - thus functioning as regular non-technical users to the project instead of researchers. However, as they are indeed researchers, the feedback they may provide on the project outcome could be so much more valuable.

- What archaeologists care about in the end is the data plottings, and Data Centres will not provide such processing. So what PEG did was implement specific software capabilities that will address the needs of non-technical data users (i.e. archaeologists), as to allow them to search for the plots or false-colour graphics they need. This piece of middleware is a conceptual key feature of the project in terms of deliverables.

- Climate data is usually archived in binary format, so it's often not easy to process. UK Met Office provided lots of info, often incomplete or in old formats. The adaption process of raw data to the project needs was very interesting and worth disseminating.

- Climate models were written in FORTRAN. When re-written or translated into C++, the results would vary for the same data arrays due to specific treatment by the code. That poses a quite amazing challenge in terms of model interpretation.

- When asked on whether researchers provided enriched metadata for their data, the answer was there's usually an input in terms of past experiments, i.e. "this is the data outcome of such and such experiment when changing initial conditions in such a way". Such-and-such experiment would be described the same way until one was reached that wasn't described at all.

The fact that none of these project aspects is recorded or discussed at the project blog poses a question on whether an external approach to data management projects might collect and disseminate very interesting information that researchers may not consider relevant enough to discuss from project blogs. Such an external approach to running projects might be carried out by data librarians in order to
share these specific project details with the data management community.

For whatever it may be worth, Sonex would be keen to do this kind of job for the MRD community.

Tuesday, 5 April 2011

I2S2 Project workshop at RAL-STFC


  Along a busy week in terms of research data management events (due to be shortly reported from this blog), last Friday Apr 1st Sonex had the opportunity -thanks to Simon Hodson, JISC MRD programme manager- to attend the I2S2 Project workshop at the Rutherford-Appleton Laboratory (RAL) at STFC in Didcot. I2S2 -standing for 'Infrastructure for Integration in Structural Sciences' is a JISC MRD project ending in Mar 2011 aiming to "identify requirements for a data-driven research infrastructure in "Structural Science", focusing on the domain of Chemistry, but with a view towards inter-disciplinary application".


Several presentations were delivered along the meeting: Brian Matthews on the I2S2 project achievements, ICAT architecture and CSMD metadata standard, Brian McMahon, International Union of Crystallography (IUCr) on 'Information Management and Publication in Crystallography', Tom Griffin on TopCAT GUI for management of data coming out of STFC ISIS and DIAMOND facilities, Steve Androulakis on the TARDIS ANDS-supported project at Monash University, Mark Borkum on OreCHEM files, Chris Morris on on PiMS (Protein Information Management System) and Juan Bicarregui on the EU PANData project.

Along the IUCr presentation the need was identified for filing & preserving different data categories such as raw measurements, processed numerical data, derived info and the paremeters. The convenience of providing access to raw diffraction images was also stressed along the talk, these files being a few GB in size, and thus not large enough for Data Centres but too big for sites such as CCDC. A review on Crystallographic Information Framework (CIF) file formats was provided, with imgCIF being used for raw data storing out of the experiment, .fcf for including structure factors after data reduction and a final stage of structure solution and refinement being performed in the lab before the author starts formatting those into a IUCr paper, which would translate CIF into SGML for producing final fcf, cif, pdf and html versions.

Raw data was mentioned to be kept for 183 days at SFTC and 3 months at Australian Synchrotron (in which TARDIS is involved), and a discussion followed on the fact that some agreement shoud be reached on the kind of data that ought to be stored and preserved. The process of attachment of DOIs to datasets was also discussed, IUCr being presently involved in projects such as XYZ or Open Bibliography in order to promote this objective.


A TopCAT demo was provided by Tom Griffin. This open source GUI (see image above) is being used for storing raw data from STFC facilities such as ISIS and DIAMOND. TopCAT provides access to its contents through an open registration system, thus operating as a sort of STFC institutional data repository, and would be potentially applicable to other institutions, facilities and disciplines.

TARDIS presentation by Steve Androulakis, Monash Univ, Australia, mentioned their using of XML/METS metadata standards for research data description at the federated institutional repository-platform initially meant to store X-ray diffraction images, later evolving into a much larger initiative with application into microscopy (MicroTARDIS), particle physics and gene processing through the Squirrel software.

Finally, extra presentations were delivered on PiMS (Protein Information Management System) by Chris Morris, STFC and on the European PANData project by Juan Bicarregui, STFC e-Science. PANData aims to build Photon and Neutron Data Infrastructure through a consortium of European synchrotron facilities and neutron sources.


A final summary was made on the whole set of presented I2S2-related features (imgCIF, CIF, IuCr/XML/RDF BIBLIO, PDBML, CML, ICAT, TopCAT, ICAT Lite/CSMD, TARDIS, PiMS, PANData, NeXuS) by mapping them on the I2S2 Idealized Scientific Research Activity Lifecycle Model (see image above - may click on it for an updated version). References were also made to other initiatives not represented at the meeting such as Quixote Project for Computational Chemistry CML data management or Protein Production and Crystallization.

Sunday, 13 March 2011

Strategies for research data deposit in ongoing data management projects


  Prior to start performing pattern analysis for research data deposit into (institutional or subject-based) data repositories –whether or not open access– first step by Sonex is to scope ongoing projects dealing with that kind of deposit, as well as already closed projects which supplied relevant guidelines on the subject. A list of projects working on data management follows, with their specific approach on how to deal with actual data deposit as taken from project blogs:


TARDIS (Monash University–Australian National Data Service).
“There is a pressing need for the archival and curation of raw X-ray diffraction data. However, the relatively large size of these datasets has presented challenges for storage in a single worldwide repository. This problem can be avoided by using a federated approach, where each institution or university utilizes its institutional repository”.


ADMIRAL: A JISC-funded data management infrastructure for research across the life sciences.
"The purpose of the ADMIRAL Project is to create a two-tier federated data management infrastructure for use by life science researchers, that will provide services (a) to meet their local data management needs for the collection, digital organization, metadata annotation and controlled sharing of biological datasets; and (b) to provide an easy and secure route for archiving annotated datasets to an institutional repository, The Oxford University Data Store, for long-term preservation and access, complete with assigned Digital Object Identifiers and Creative Commons open access licences".
(See Oxford University Library Services' Databank)


XYZ Project. “The XYZ Project will create a demonstrator of a new workflow for publishing data in support of full-text. The author prepares data for publication (if possible with validation) in a third-party trusted repository before the paper is submitted to a publisher. Our software will manage the deposition, release to reviewers, dis-embargo and for conventional publication or as a data journal. Two Open Access publishers (International Union of Crystallography and BioMed Central) are engaged with the project and will test the new workflow”.
Anticipated Outputs and Outcomes: A demonstrator repository hosted by the IUCr.


FISHnet: Freshwater information sharing network. “This project will allow researchers in multiple academic, governmental and voluntary-sector institutions to share their data. Data will be held securely in a sustainable subject repository which preserves and disseminates multiple datasets as part of the FreshwaterLife.org information portal. Data creators will be able to manage access rights to their content, from Open Access to sharing with trusted colleagues”.


DMBI: Data Management in Bio-Imaging. “The quantity of data generated by modern high-throughput bio-imaging systems presents a significant challenge in both data management and processing. Furthermore, there is no explicit system/way to record the processing algorithms and parameters that are used to produce results. Thus there is no strong link between images, software and results. This projects aims to address these issues”.
Anticipated Outputs and Outcomes: Build a prototype DMBI system around OMERO.


CaiRO: Curating Artistic Research Output. “No prominent subject-based repository exists to act as the custodians of arts practice-as-research data. Where institution provision for data management is in place (for instance, an institutional repository service) the arts researcher-practitioner cannot always rely on an understanding of the special nature of arts research data. More commonly, data is retained in departmental collections, built and maintained by small teams which often include researchers themselves”.


BRIL: Biophysical Repositories in the Lab. “The BRIL project aims to enhance the repository facilities at the Randall Division of Cell and Molecular Biophysics at King’s College London. This will involve:
» Embedding the repository within the researchers’ day-to-day research and experimental practices;
» Integrating the repository into the wider King’s infrastructure”.
Example of KCL “internal” repository: Mutation Testing Repository.


ADS+: Enhancing and Sustaining the Archaeology Data Service digital repository. The project aims to “Increase the sustainability of the ADS, by implementing Fedora (Flexible Extensible Digital Object Repository Architecture). This is a world-leading open source digital repository application which will allow the automation of many ADS curatorial functions, according to the Open Archival Information System (OAIS) Reference Model (ISO 14721:2003). This will help ensure the long term preservation of all ADS digital archives, as well as making the ADS archival procedures more cost-effective”.


IDMB (Institutional Data Management Blueprint) Project, U. Southampton.
The project’s aims are to provide the University of Southampton with a ten-year roadmap for delivery of a comprehensive data management infrastructure.

[IDMB Recommendations] The data management audit and gap analysis indicates where improvements can be made in the short, medium and long-term to improve data management practices and capabilities at the University. The following preliminary recommendations are put forward for short (one year), medium (one to three years), long (more than three years) term action.
[Short Term (1 year)] Crucial to supporting researchers is the consolidation of data management into a coherent framework that is easy to understand, use, and has a sustainable business model behind it. A number of major recommendations are put forward here for the short-term:
Create an institutional data repository
• Develop a scalable business model
• One-stop shop for data management advice and guidance


MaDAM: Pilot data management infrastructure for biomedical researchers at University of Manchester.
A pilot infrastructure for Biomedical Researchers at the University of Manchester, which covers data capture, data storage and data curation. This infrastructure comprises procedural support, hardware and software.
[18/03/2010] The development team have built a prototype data management front end which fits a generic set of needs amongst our Life Sciences researchers. It is aimed at being flexible enough to allow researchers themselves to assign attributes (i.e. metadata) to their experiments and datasets for them to be usefully categorised and tagged. The prototype is also entirely dispensable and intended as a catalyst for feedback from our use cases on their specific functionality requirements.


DISC-UK DataShare Project. The DISC-UK DataShare project, led by EDINA National Data Centre and the Edinburgh University Data Library, with partners at the Universities of Southampton and Oxford, has advanced the current provision of repository services for accommodating datasets in the UK.
Key conclusions: 1) Data management motivation is a better bottom-up driver for researchers than data sharing but is not sufficient to create culture change, 2) Data librarians, data managers and data scientists can help bridge communication between repository managers & researchers, 3) Institutional repositories can improve impact of sharing data over the internet.

Thursday, 3 March 2011

Repository take-up and embedding: the future of repositories


  Being already in Birmingham for the JISC Deposit Project Meeting on Mar 1st, Sonex stayed in town for attending the JISC Repositories Take-Up and Embedding Meeting as well. Start up meeting for this new JISC programme aimed to outline the future of repositories, dealing with specific issues such as (automated) deposit, shared services like RoMEO or OpenDOAR, repository integration into general software infrastructures for research information managament and promoting national (via RSP) and international (via KE, COAR and OpenAIRE) collaboration.

Six projects were presented along this programme start up meeting:

- Bringing a Buzz to NECTAR (Miggie Pickton, University of Northampton)
- Hydrangea: letting the repository flower (Richard Green, University of Hull)
- MIRAGE 2011: Repository Enrichment from Archiving to Creation (Xiaohong Gao, Middlesex University)
- Enhanced interface design for supporting take-up and embedding of the Glasgow School of Art research repository, including visual
engagement with practice led and applied outputs (Robin Burgess, Glasgow School of Art)
- eNova (Marie-Therese Gramstadt, VADS)
- EXPLORER: Embedding eXisting & Propriatary Learning in an Open-source Repository to Evolve new Resources (Alan Cope, De Montfort University)

An extra postprandial presentation on repository consolidation within a university research information management environment and the way it was done at University of Glasgow Enlighten IR was delivered by Willian Nixon. Statements like "Silos are the past, embedding repositories -through the use of tools like Sword or LDAP- is the future" made the point on how repositories should evolve in the future. According to William, repositories are to exploit new opportunities for data mining, business, intelligence, KPIs, analytics, 'stickiness' and visibility (some of these issues being thoroughly dealt with at Enlighten repository blog).

There was a remarkable presence of image-related projects among the presentations, Glasgow School of Arts, eNova and MIRAGE 2011 dealing with archiving of images into repositories one way or another. This is great news for momentum-gaining development of new information infrastructures in the area (also traceable at the JISC Deposit Programme meeting the day before), which will no doubt benefit from these projects outcomes.

After watching project presentations from a Sonex point of view, it seems they could particularly benefit from interacting with JISC Deposit projects in terms of implementing resulting strategies for automated content ingest into repositories. A handful of the take-up and embedding projects would thus be the soundest candidates for initial "customer implementation" of the various resulting methods for quick population of repositories with institutional research output (the take-up bit, prior to embedding) coming from the Deposit strand. As these projects will run
until the end of 2011 and the ones from Deposit strand should deliver around July, interaction among them could probably be easily achieved.

There was one particular project among those presented that captured Sonex's attention: MIRAGE 2011, Middlesex Medical Image Repository with a Content-Based Image Retrieval Systems Archiving Environment. MIRAGE is both an image-related repository project (as it deals with medical images) and a research data project, and it's this latter feature what gets it fully within scope of Sonex activity with regard to research data management. Ongoing data management projects (either JISC-funded or otherwise) usually deal with either numerical or textual data, but projects dealing with the deposit of graphical research data are rare (save for Data Management in Bio-Imaging - DMBI project run at The John Innes Centre, BBSRC, Norwich).

A couple of references were shared with MIRAGE project manager Dr. Xiaohong Gao, 'Feeding Neuroimaging Repositories' poster presented at OR2010 Madrid last July by a team of Universitat Autònoma de Barcelona (UAB)-Hospital de la Santa Creu i Sant Pau researchers in Barcelona, and the MIDAS/National Alliance for Medical Image Computing (NAMIC) medical image repository as to promote synergies among different projects on the same area.

The meeting presentations will shortly be available.

Wednesday, 2 March 2011

JISC Repository Deposit Programme Meeting in Birmingham


  A JISC Repository Deposit Programme meeting was held on Mar 1st, 2011 at Maple House Birmingham. Under coordination from Balviar Notay, JISC manager for the Deposit projects, presentations were delivered from representatives of the four presently running projects under JISC Deposit call: DepositMO (Steve Hitchcock, U Southampton), DURA (John Norman, UCam), RePosit (Ian Tilsed, Leeds U) and Kultivate (Marie Therese Gramstadt, VADS). Additional presentations were done for the deposit-related Open Access Repository Repository Junction (OA-RJ) project (Theo Andrew, EDINA), Sword v2 (Richard Jones - Symplectic) and Sonex (Pablo de Castro, Carlos III University Madrid) projects.


Lots of interesting issues were raised and discussed along the set of presentations, and specific teamworking activities were later carried out for promoting cooperation between projects. This was the first opportunity for representatives of all projects involved in the JISC Deposit programme to personally meet the other projects and learn about their progress and potentially complementary findings.

Several complementary visions of deposit were outlined along the workshop: a quite technical one from projects such as DepositMO and Sword, an advocacy-focused approach from RePosit project aiming to increase engagement to repository and a vision of repositories as potential suppliers of the global institutional research output required for REF purposes from DURA.

Steve Hitchcock (DepositMO, implementing Sonex usecase scenario nr 4, Deposit via personal software) delivered a few demo examples of Swordv2-assisted deposit into the DepositMO test repository via local computer file manager, including deposit of previously parsed full-text document ingesting metadata as well and achieving the metadata+object transfer. A key question on document deposit for management vs publishing purposes was also raised along DepositMO presentation: are repositories (or could they evolve into) a proper environment for document management or does the Open Access philosophy prevent them from being used as cooperative tools for example for pre-print edition by a group of authors?

DURA and RePosit projects, implementing Sonex usecase nr 2, CRIS/IR integration, are both dealing with making deposit as easy as possible for the author community by ingesting previoulsy synced inputs from Mendeley and Symplectic Elements into IRs (DURA) and specificallly “increasing engagement with repository” (RePosit) by designing a set of awareness-raising materials and campaigns later to be shared with other projects.

Kultivate, aiming to increase deposit in the arts and design environment, is both the newest and possibly the most innovative project in the strand. Repository development having been strongly focused on research papers as a main research output, work on so far underexploited creative arts materials gives Kultivate the opportunity to set new standards and provide new resources to the Open Access repository community.


Further presentations for projects providing general-purpose deposit infrastructure followed, such as EDINA Open Access Repository Junction (OA-RJ) middleware for discovery and Sword-assisted deposit. OA-RJ is already live-testing its broker for automated transfer of publisher or subject repository content inputs into specific target repositories. Richard Jones described the ongoing process for developing Sword-v2, which will deliver fine-tuned functionalities for metadata+object automated transfer to the rest of the Deposit projects and the wider repository community, resulting in higher deposit rates. Finally, a Sonex presentation stressed the need for re-examining Sonex deposit usecase scenarios for covering new types of materials such as research data, creative arts materials and learning materials. Sonex also suggested common strategy for measuring success of JISC-funded deposit projects being designed at Birmingham City University Evidence Base might include specific questions to be asked to repository managers such as whether any given automated deposit strategy was used for content ingest purposes besides specific strategies for measuring success devised by projects themselves.

The workshop presentations will shortly be available at the Deposit wiki. Once Deposit projects are completed another programme meeting will be held for sharing conclusions and examine case studies and success stories as to widely implement resulting solutions.

Sunday, 16 January 2011

"On such a full sea are we now afloat"

Such quotation -from W. Shakespeare's 'Julius Caesar'- closed Drs. Eefke Smit's talk "Taking the Current when it Serves: Research Data from the Publisher's Perspective" she delivered along 'Academic Publishing in Europe': the APE 2011 conference, held at the Berlin-Brandenburg Academy of Sciences in Berlin, Jan 11-12th, 2011.


Aiming to gather some facts for its ongoing analysis on research data management and its deposit into repositories, Sonex just attended APE2011, a meeting for the publishing industry and its environment held yearly in Berlin since 2006. The conference organisers do regularly publish a brief official report shortly after the event celebration (reports on previous APE editions
available here, report on this edition due shortly).

This particular visit to Berlin offered the chance to attend yet another event besides APE2011: the SOAP Symposium. Final report by the SOAP (Study of Open Access Publishing) project survey was presented along this one-day meeting, held on Jan 13th in the Goethe Room of the renowned Harnack-Haus in Berlin. The SOAP project describes and analyses the open access publishing landscape as well as exploring the risks and opportunities of the transition to open access publishing for libraries, publishers and funding agencies - see preliminary survey results, final report will be available as of next March.

The conference programme for APE2011, entitled "Smarter Publishing in the New Decade", included promising topics such as evolution of peer-review and ways to improve it, the so-called data deluge, business opportunities in China and how Open Access is becoming increasingly mainstream within the publishing environment. Discussions on those matters were lively both at round tables and at lunch pauses. Sonex interest being mainly on research data management, this report will subsequently focus on presentations and debates on the subject.

On Tuesday Jan 11th afternoon, a session was held on “The Data Deluge: to Drown or to Swim?”, chaired by Bob M. Campbell. Herbert Gruttenmaier, INIST-CNRS, started his presentation "Helping to Ride: a look at data sharing and access policies" by reminding that, since we were in Berlin, the definition of an Open Access Contribution on page 1 of the Berlin Declaration on Open Access to Knowledge includes “raw data and metadata”. Some highlights from his talk were:


  • There is a large number of Data Sharing Policies being defined by administrations, institutions, funding agencies and publishers themselves under the guideline "data should be made as freely and widely available as possible". See for instance NSF’s requirement for submission of data management plans of May 10th, 2010, under general policy statement “Investigators are expected to share with other researchers, at no more than incremental cost and within a reasonable time, the primary data, samples, physical collections and other supporting materials created or gathered in the course of work under NSF grants”.
    Or the very recent (Jan 10th, 2011) commitment by a group of major international funders of public health research to “work together to increase the availability of data emerging from our funded research, in order to accelerate advances in public health”.

  • Publishers such as BioMed Central were featured as high-profile supporters of Open Data (see Dec 11th, 2010 post at this blog), and NPG editorial policy on dataset sharing was specifically mentioned along the talk, as well as the Brussels Declaration on STM Publishing statement that “Raw research data should be made freely available to all researchers”. Finally, discipline-based data policies such as PaN-Data Scientific data Policy Draft for Scientific Data Management Framework at European Photon and Neutron Facilities or the Joint Data Archiving Policy (JDAP) adopted in a coordinated fashion by Dryad partner journals.

  • Not everything is that simple though: the Nov 2009 "Patterns of information use and exchange: case studies of researchers in the life sciences” RIN report shows that researchers are not so eager to share their data with others, and that ‘one-size-fits-all’ information and data sharing policies may not achieve the goals there are aiming for, namely scientifically productive and cost-efficient information use in life sciences.

Drs. Eefke Smit, International Association of STM Publishers, provided a counterexample for these growing data sharing policies by publishers along her talk on "Research Data from the Publisher's Perspective" by describing the Journal of Neuroscience policy of no longer taking supplementary material from authors since Nov 1st, 2010, the procedure posing too heavy a burden on paper reviewers.
She also warned of the so-called data deluge, according to which tera- and petabite sized datasets will increase their share in research projects in upcoming years.
However, when researchers are asked where they would like to submit their research data, the answer is more often than not "publishers". This brings along the issue of research data preservation: results of an internal survey by STM Publishers show what she called “an improvable situation” with regard to preservation.

Planned talk “Data Publishing in the Context of the ICSU World Data System” by Dr. Michael Diepenbroek, Director of WDC-MARE/PANGAEA, University of Bremen, went finally off the conference programme. However, the next speaker, Dr. Jan Brasse, Managing Director of DataCite, provided some information on the progress of one of the main databases for research data in the geosciences area, by for instance stating there was “a wide cooperation between Elsevier and PANGAEA via DOI-based external links from online papers” at the former’s platforms. This kind of cooperation between publishers and international databases for handling research data might be useful for tacking the abovementioned data preservation issues.
Dr. Brasse, affiliated with the German National Library of Science and Technology Hannover, described as well the evolution of the DataCite international project as it gets carried out by local member institutions: as of Dec’10, over 1M records are already registered with DOI names at datacite.org. Perspectives for the project include setting up of a Central Metadata Base as of Jun'11; DataCite becoming a harvest point for third parties such as WoS; and cooperation via CrossRef for data-article lookup.

The data management session ended with the talk on “Managing Publication and Research Data: the eSciDoc Research Infrastructure” by Dr. Malte Dreyer from Max Planck Digital Library (MPDL). eSciDoc is as a joint project of the Max Planck Society and FIZ Karlsruhe, funded by the Federal Ministry of Education and Research (BMBF), with the aim to realize a next-generation platform for communication and publication in research organization. Further eSciDoc projects mentioned along the presentation and dealing with research data management were ‘Astronomer‘s Workbench’ (astronomy), Lifecycle Logger (biochemistry) and BW-eSci(T) for computational linguistics. DARIAH (Digital Research Infrastructure for the Arts and Humanities) –in whose development eSciDoc is directly involved- and CLARIN (Common Language Resources and Technology Infrastructure) projects were repeatedly highlighted along the session as leading EU projects on development of digital research infrastructure (including data management) for the Humanities and Social Sciences.


A joint panel discussion was then held after the presentations on research data management, with speakers taking questions from the floor. Alicia Wise, Elsevier Director of Universal Access and former archaeologist raised the issue of costs attached to research data management and who should fund them: it was agreed by the panellists that national funding bodies should assume the cost of data management. Along her question Dr. Wise incidentally mentioned that data management at the archaeological research project she used to work for succeeded only thanks to researchers dedicating 50% of their time to data curation. This aspect of dataset deposit will be examined by Sonex in order to identify alternative (automatic) curation procedures currently being used to relieve researchers of the data curation burden.

The data management issues extended well outside the session specifically devoted to them and into the Innovation session held next day, where Portland Press Adam Marshall presentation on the Semantic Biochemical Journal and Project Utopia at the Manchester School of Computer Science did extensively deal with data handling (see “Calling International Rescue: knowledge lost in literature and data landslide!” at Biochem J. (2009) 424, 317–333 for a review on “how to provide new ways of interacting with the literature, and new and more powerful tools to access and extract the knowledge sequestered within it”).

At the end of the data session panel discussion Dr. Eefke Smit synthesized the three challenges of research data management: normalization, standardization and migration. She did also remind the audience of verses following the one quoted in the title of this post:

(…) On such a full sea are we now afloat,
And we must take the current when it serves,
Or lose our ventures
.