Metadata fatigue is not a fatality

Just as Bill Murray relived Groundhog Day in the famous 1993 movie, researchers may feel trapped in a time loop when forced to re-enter the same metadata into countless formats. Again and again. Software Heritage asked Daniel Garijo, its new ambassador, how to break the metadata fatigue curse without compromising on quality.
Daniel Garijo is an associate professor at the Artificial Intelligence Department of the Computer Science Faculty of Universidad Politécnica de Madrid and part of the SciCodes consortium, a group of research software registries and repositories heavily invested in improving Open Science practices. Garijo first heard of Software Heritage when collaborating in the Research Data Alliance. His expertise lies at the intersection of knowledge representation and open science, with a particular interest in research software and reproducibility in academic works. Research software is a key asset here, as it supports the results described in academic publications. Improving the adoption of good practices, which include FAIR Research Software metadata, is one of his goals.
As an ambassador, Garijo aims to raise awareness of features like Zenodo’s automated archive and to foster community-driven improvements to software metadata representation and adoption through the CodeMeta initiative, a community standard to represent research software metadata. To ensure uniformity and consistency in descriptive metadata, Software Heritage decided early on to use the CodeMeta vocabulary internally when indexing intrinsic metadata.
Created in 2015 from the FORCE11 Software Citation Working Group, CodeMeta provides a shared, machine-readable vocabulary for describing research software. Today, it is used or supported by a growing number of infrastructures and organisations, including swMath, Zenodo/InvenioRDM, HAL, ascl.net, and many scientific registries via the SciCodes consortium. CodeMeta enhances research software descriptions via semantic metadata.
Learn more:
- Gruenpeter, M. (2025, November 25). CodeMeta: Metadata, Ontologies & Interoperability for Research Software. Zenodo. https://doi.org/10.5281/zenodo.17713873
- Deep Dive into the archival of Software Metadata
How did you get interested in software metadata? Why did you start exploring this topic?
I started exploring research software metadata, or scientific software metadata, as we also called it, because I got involved in the MINT Model Integration project as part of the Defense Advanced Research Projects Agency (DARPA World Modelers program. Our project collaborators and I had to create a registry of scientific model metadata, which has many similarities with research software itself. In fact, we represented it as an extension of research software.
Also, my interest in metadata grew through another project calledOntoSoft, which was funded by the National Science Foundation under the EarthCube Initiative. We described software for scientists in geosciences.
Both projects significantly familiarized me with research software metadata, leading me to rethink how to represent different software aspects to improve actionability and reproducibility, rather than focusing solely on findability.
In particular, within the OntoSoft program, I helped translate the questions captured as metadata fields into terminology more familiar to researchers. Then I started getting involved in CodeMeta.
In your opinion, what are the three main challenges currently impeding better software metadata quality?
Challenge #1: Too many schemas and too many formats
In my view, the first challenge is that there are many, many, many different schemas, which is one of the issues that CodeMeta aims to address. By having such a wealth of different fields and different standards that people can follow, it’s sometimes very difficult to have everyone provide metadata in a homogeneous format.
Now, there are tools that will help extract and homogenize the information. And of course, agents and large language models may play a role. But there is also a risk that this metadata could be more diverse or even hallucinated in certain parts.
So the challenge is the number of vocabularies that exist to represent software metadata, which I think we can help address with CodeMeta.
Challenge #2: Be kind, do not make me rewind
The second challenge impeding better software metadata quality is the burden we place on researchers to provide the same information repeatedly, particularly when a new format introduces a new restriction, such as using normalized identifiers for persons or organizations.
Researchers care about their software’s metadata. However, their concern is the burden of entering structured information into yet another format. When I speak with them, they invariably respond with examples like: « Oh, yes, I’ve already included all this requested information in the report I sent to my funder, » or « You may find some of this data in my README file, » or « I’ve written extensive documentation”.Their question is often: « If you need this in your platform’s specific format, why can’t you extract it yourself? »
These practices place a burden on researchers, yet they are tasks that could be automated. And I think that is a very fair concern.
Challenge #3: One metadata field and so many different approaches
The third challenge is usually one that I experience a lot, which is the social aspect: getting consensus on which information is important to get for performing certain tasks. For example, platform managers’ top priority is improving findability. For them, having keywords is important. Having descriptions matters. Having a title is crucial. Having the authors properly listed is a prerequisite to attribute credit.
Now, let’s consider the researchers’ metadata approach. While researchers certainly care about attribution, their primary concern is ensuring reusability for others (and for themselves). Sometimes, you may hear: “I don’t care that much right now about giving credit, because I’m just reusing their software to see if It’s viable for my problem. But what if I can’t access the executable, if I can’t know which data should be used with? .”
Sometimes there are overlaps between these views. The view of:
- The one who tries to find records,
- The one who needs to reuse software, and
- The one who needs to put all these things over there to get the credit.
But sometimes they are disjoint. And trying to standardize the output so that everyone provides a homogeneous representation is also a challenge. And that is the challenge that we face when people ask to introduce new properties in existing schemas.
How do intrinsic and extrinsic metadata relate to each other? How to ensure metadata consistency across all these files?
I think this is the main goal of CodeMeta. CodeMeta originated from the need to map every single approach together, given the numerous existing mismatches. However, I think there is now an opportunity to have an interoperable standard for research software metadata.
As for metadata consistency, my colleagues and I have been working in my lab on this problem by developing the tool RSMetaCheck. This tool detects common metadata quality issues (pitfalls & Warnings) in software repositories. RSMetaCheck analyzes SoMEF (Software Metadata Extraction Framework) output files to identify various problems in repository metadata files such as codemeta.json, package.json, setup.py, DESCRIPTION, and others.
For example, what happens if you see two different published dates? Well, you can dig into the GitHub history and check whether the date of the first release is correct. But you can also automate this task.
You may also need LLMs for more complex tasks. For example, you might be using a REUSE specification for different license files in your project, but sometimes there are licenses listed with no context. What is the source of truth? What applies where? I believe that having such tools in the future may help us by raising issues, similar to how Dependabot addresses security concerns. This will help polish and homogenize metadata quality. But I think that so far, people have been concerned about software quality, not about software metadata quality that much.
How would you define CodeMeta in a single sentence?
CodeMeta is two things at once. First, CodeMeta is an interoperable metadata schema for research software metadata. Second, CodeMeta is a set of crosswalks, which capture the effort put in by the community to map all existing software metadata schemas into a single, well-defined, or common representation.
In one sentence: CodeMeta is a harmonization effort to bring together research software metadata.
Most of the CodeMeta fields are optional. What are the three fields that can have a great impact for researchers if properly filled?
Every time I must complete a new CodeMeta file, I adopt the perspective of the software author. Therefore, I would put my effort into the following fields:
- The name,
- The description,
- The citation.
Because I want people to find my software, I want people to know very easily what my software does. These are the main things that people look at to decide whether the software is suitable for their purposes or not. And finally, the citation field, because this way, I get the credit. It’s the fastest path for me to get credit for my research software.
What advice do you have for people getting started with CodeMeta?
Go for a minimum metadata file that is good enough for your purposes, and if needed, then start making it more complex.
Many infrastructure managers will argue: “You have to use these identifiers, and then you have to use these proper objects and URLs.” And that sometimes overburdens users.
Describe who you are, what this software is for, how to start it, and how to get credit. Focus on that. And once it’s done, if you really want to improve your findability, add more keywords. Tell me the application domains. If you want to improve your usability, then tell me what the dependencies are. Tell me where the README is, how to get started, and how to install it. And if you really want to have a proper linkage across platforms, then, well, don’t use names; use ORCIDs, right? Or use the roles of your organization. Thus, the transitive credit will be propagated a little bit more easily. And that’s the way to get more impact. But don’t start everything at once because otherwise you’ll get overwhelmed.
My recommendation is to use existing tools that are there to help you, like the AutoCodeMetaGenerator or the CodeMetaGenerator. AutoCodeMetaGenerator is actually pretty good if you already have some metadata in your repository.
Alternatively, you can try to use a large language model to help you complete the descriptive files, but always check the results. Regardless of the tool used, always verify that the content matches your intent. It’s better to have a short, well-described, concise metadata file than a very long metadata file with many inaccuracies.
Software Heritage Ambassadors are volunteers who offer expert advice in various sectors and languages on how to use our services. Here’s more information on how to book one for a free consultation.
If you’d like to connect with Daniel Garijo, please reach out using this email: daniel.garijo[@]upm.es
We’re also seeking passionate individuals and organizations to volunteer as ambassadors and help grow the Software Heritage community. If you’re interested in becoming an Ambassador, please share a bit about yourself and your connection to the Software Heritage mission.

