Submind YouTube summaries
Thumbnail for OSF Essentials: How to Use PIDS on the OSF

OSF Essentials: How to Use PIDS on the OSF

Watch on YouTube

Video summary

This webinar from the Center for Open Science introduces Persistent Identifiers (PIDs) as essential tools for enhancing research visibility and data management on the OSF platform. PIDs are defined globally unique, persistent, interoperable, and machine-readable identifiers that ensure resources like people, organizations, and digital objects can be reliably located over time regardless of URL changes. The session explains how these identifiers disambiguate between entities with similar names, such as distinguishing one researcher from another, and serve as stable links to publications known as Uniform Resource Identifiers (URIs). Unlike standard URLs which may break or redirect unpredictably, PIDs are maintained by organizations committed to resolving them consistently, thereby ensuring long-term accessibility for research outputs. The core benefit of using PIDs lies in their ability to construct an information graph rather than relying on isolated records. By linking identifiers such as ORCID for authors, ROR for institutions, and DOIs for publications or datasets, the OSF can infer complex relationships between researchers, organizations, and works automatically. This interconnected network allows users to answer sophisticated questions about research impact by aggregating data across multiple repositories without duplicating information manually. For instance, if a preprint receives a DOI registered with Crossref while being authored by someone linked via ORCID at an institution verified through ROR, the system can seamlessly associate all these elements into a cohesive profile that expands as new works are published or versions are updated. To leverage this infrastructure effectively, researchers are encouraged to register for their own ORCID iDs and link them directly to their OSF accounts using single sign-on verification. Once connected, any DOI generated for preprints on the platform will automatically populate the researcher's profile without additional administrative effort. Furthermore, users can enrich project registrations by attaching DOIs or other PIDs to supplementary materials like code repositories, data sets, and published papers through a simple interface that displays colored icons upon successful linking. This practice not only reduces the burden of metadata entry but also ensures that research outputs are precisely identified for publishers and aggregators who rely on these standards for dissemination and citation tracking. The webinar concludes by highlighting how PIDs have become a mandatory requirement in major funding policies, including those from the National Institutes of Health and UK Research and Innovation, underscoring their critical role in modern open science practices. Looking ahead, future sessions will cover digital preservation strategies to ensure data longevity, registration templates for structured project design, and advanced usage of the OSF API. Additionally, new self-paced training modules on topics like pre-registration and data management are available for researchers seeking to build practical skills throughout their research lifecycle. By adopting these identifier systems now, scholars can contribute to a larger global network that makes it easier to discover, cite, and connect valuable research contributions across disciplines and institutions.
Read the full video transcript
Hello everyone, welcome to our webinar today. Um, as folks are coming in, I will just welcome you. Uh, this is the latest in our 2026 series of OSF Essentials webinars. Uh, these are focus sessions on different topics each month built around the questions, workflows, and practical needs that users most often want help with. Uh, so join us on the last Thursday of every month. We have three or four more of these scheduled for this year. Um, this month we're going to talk about persistent identifiers, otherwise known as PIDs, and how to use them on the OSF. Um, I'll have some more information as well at the end of our session about our future series of webinars and other training opportunities from Center for Open Science. So, just some reminders before we get started. Um, please use the Q&A box for questions. This should be towards the bottom of your screen, probably in your Zoom controls. Um, it says Q&A. Um, type your questions in there. It will be much easier for us to track them and make sure that we get them answered. Um, you can use this the chat, however, to introduce yourself or or make comments. Um, but we do end up missing questions if they end up being in the chat rather than the Q&A. And just as a reminder, you will receive an email with the recording of this session and the slides um probably early next week. So, um, don't worry if there are URLs that show that you don't get to copy down in time. You will receive the slides with all those links. So, let's get started. Today, we're specifically going to talk about persistent identifiers or PIDs and what they are. Then, we'll talk about how PIDs improve research. And finally, we'll talk about using PIDs on the OSF. Okay. So, what is a persistent identifier? Uh, persistent identifiers, PIDs, are globally unique, persistent, interoperable, and machine readable identifiers. So, essentially, they're unique numbers assigned to people, places, or things by an organization. But it's not enough to just be unique. That organ organization that creates the number has to have a guarantee that they will maintain it. those organizations register or mint the identifiers and then keep a consistent record for them and if it's a web- based resource make uh also the URL where they can be found so that moving forward you can always find that resource because the URL gets updated by the organization in repositories we're primarily talking about people who create and contribute to works the organizations that are associated with them and the work objects they produce the articles, the papers, the data sets. Okay, so PIDs are unique and persistent. But what's the point? Well, sometimes the identifiers are used primarily to disambiguate or distinguish between between things. So you can see here just part of a list of people named William Smith. The identifiers here work to distinguish one William Smith from each other. When works that William Smith creates use this identifier to refer to him, we can easily create a list of all of the works associated with one William Smith over another. Another way you're probably familiar with PIDs is when they're used as links to publications like these DOIs. When PIDs are used as links, they're sometimes known as uniform resource identifiers or URIs. That's very similar to universal resource locator or URL which you may be familiar with. The difference is that the URI is a persistent link from an organization committed to resolve the link to a document in the future. Digital object identifiers or DOIs are a widely used PID for objects in the research environment. But there are also persistent identifiers for objects uh from other organizations with other names such as arcs, handles and pearls. We when we can link PIDs within a system within a repository for example we can create an infrastructure based on those persistent links. So this means that in OSF for example when we store data we don't have to store data about the same person or the same organization multiple times in multiple records possibly spelling it multiple ways or creating other inconsistencies. So let's say we have this object in OSF and it's created by this user for a project funded by this institution. Each entity, the person, object or organization in OSF has an ID, either an OSF, what we call a GID, which is unique, a unique identifier, but not a PID. And then uh they may have another persistent identifier like orchid for this person and auror ID for this institution. Talk a little bit more about those in a second. in the OSF object. Then we only have to store the identifier and not all the information. When we present this record, when we present back to you information about this object, we can fill in the correct name or other info from the source record. And then the identifier can be used over and over again in all records. So here is the same identifier for this creator in another work that they've created. When PIDs are assigned relationships to one another like this, it makes what's called an information graph. So this user has uh an orchid which is shown here and they are uh they work for this organization which has what's called a roar ID for research organization registry. This user then creates this work which has a DOI which is a digital object identifier. That DOI is in turn published in a collection volume that has its own DOI, a separate DOI. So the power of graphs then is that they can be used to infer and expand relationships. So if I know the text was written by this user and this user is affiliated with this institution, I can by inference associate the text with the institution. Creating graphs of PIDs based on relationships between identified items instead of just flat individual records allows you to answer more complex questions. So in the graph here which is the image on the right going to annotate a little bit. Um the question is I want to find all digital objects that are connected to a research object. In this case I'm interested in this one right here. So I can tell by looking at that object that it has a relationship to this one and it has a relationship to this one. But in addition, these two objects have other relationships to these other objects. So by uh putting these all together, I now know that there is an inferred relationship between this object and that and the uh original text and this other record. So this allows us to visualize the real impact of the relationship that this work has had to all of these objects as opposed to just the ones that are listed in the bibliography. For example, by aggregating records that use the same identifier from more than one repository, more information about those items can be pieced together. So if OSF shares our graph data, um that can be used by other organizations, they can combine it with their own to build a bigger and better overall graph. And since OSF is itself multiddisciplinary and multi-institutional, we can reveal relationships that might otherwise be difficult to surface. Hold on. There we go. Pardon me. I'm having a little There we go. So, the end goal is to create an even larger network of graphs that can connect more information and help uncover more types of relationships and connections that might have been difficult to detect. This is a graph of biblometric data about journals and it shows connections between titles in seemingly unrelated areas. Other graphs could show networks of researchers and their interests which could inform decisions about research support or potential collaborations. Graphs with this level of breadth and scale are possible through PIDs and networks of shared information. PIDs are so important and powerful in fact that they're the number one item on the National Institute of Health in the United States list of desirable characteristics for data repositories. Um they're also highly ranked on the National Science and Technology Council's list of desirable characteristics. PIDs also appear in the UK research and in innovation UKRI open access policy plan S requirements and appears in the core trust seal requirements. You may not know what those organizations are, but they're basically organizations that verify that the work we do in repositories to make sure that your research work is disseminated and uh persists. Uh those are organizations that ensure that that happens. NIH specifically requires that a repository assigns data sets a citable unique persistent identifier such as a digital object identifier to support data discovery, reporting, and research assessment. The identifier points to a persistent landing page that remains accessible even if the data set is de accessioned or no longer available. So in OSF we uh primarily use three types of PIDs. PIDs for three different things. We use what's called orchid for users, roar for organizations and for funding agencies and DOIs coming from data site and crossref for um OSF objects. Crossref is an organization that registers DOIs for preprints and other written works while data set uh data site registers DOIs for other types of content like data sets. If you're comfortable with code, this slide shows how these identifiers actually get into the metadata that we share in objects in OSF. If you aren't though, don't worry. I'm going to talk us through it. So, in this top example, I've highlighted, you can see the affiliation identifier, which is a URI containing roar.org right here. And we can see that it's labeling the value University of Oxford. So, now we have the PID and the name of this organization. Later in the record, we also see how PIDs can help build relationships. In these three fields, I'm referencing three other related items to the work that we're describing here. [snorts] So, one is a version and it resolves to a DOI. Um, this is missing the prefix for the DOI, but that's how DOIs are are format formatted. Next, we have a supplement that resolves to another OSF object. And finally, this object refers to a GitHub repository. So, it's really simple, but this is how the magic happens. um we create this record, we pass it along to say data site and then they can sync up the document with the version listed here, the other document that's a version or other things from this organization or uh other records by this person. There's an orchid in this record as well which you can find if you want uh to have a little adventure. Um so uh by using these identifiers in data site they can aggregate together information on all of these uh all of the works by this person all of the works from this organization etc. So PIDs enable a graph of information in OSF but because we share metadata in other places we expand the graph into other places. So I just showed you what that looks like in term of the record that we share but let's look at an example. In this example, the user has added their Orchid ID to their account. They then uploaded a preprint to our servers and it was assigned a DOI which was registered with crossref. So now that preprint shows up in crossre and we were able to add the user's orchid ID to that metadata. This means the citation can also be scooped up into the user's orchid profile. And all the user had to do in this case was add orchid to their account. Everything else happened automatically because of the way we register PIDs for preprints and share metadata. In another example, this user has also added an orchid and they have also been affiliated on the OSF with their university through our institutional membership program. So when we registered a DOI uh for this person's project with data site, we were able to pass along the person's name but also the organization that they're a part of. So the resulting record and data site has all of that information. So these examples show how PIDs can reduce administrative burden. In these cases, one resource orchid for example benefit benefited from the data in another resource OSF. In addition to PIDs always resolving to something, if you use m machine readable data standards and PIDs like this, we can obtain the full metadata for the resource at the PID location, which can then be harvested and reused like I've shown. In an ideal scenario, a researcher administrator only needs to populate metadata in one system one time to then see it propagated down the line. So, would you like to make some PIDs? I hope so. Uh, I'm going to talk about ways you can use PIDs in OSF. You can add PIDs to your content in a couple of ways. Um, this will help to build a knowledge graph about yourself. The first is to create an Orchid for yourself if you haven't already. You can do that at orchid.org/ org/register. Once you have your Orchid, you can then allow it to autoupdate. This will grab references to things you publish that get a DOI and publish them to your Orchid record automatically. Later on, when someone looks at your Orchid record or another tool you use, pulls in information about you from Orchid, it will have all of that information. Finally, then you can associate your Orchid with your OSF account. You just choose to sign in with Orchid and use uh your same username and password you already have. You'll then go through a verification pro process and the profiles will be connected. This whole process is described in the help guide listed at the bottom of this slide and again you will receive a copy of these slides. So don't worry if this sounds like a number of steps. We have it all in a walkthrough in this guide. So now every time you create a DOI for something on OSF, if you've connected your Orchid, it will get picked up into your Orchid record. Another way you can really leverage the use of PIDs and OSF is to ensure that you've created good metadata for any object that you get a DOI for. When you create a DOI in OSF, we send a metadata record to data site or crossre like I've shown, depending on what type of material it is. They create a link and then redirect to your OSF content. They have committed to keeping those links links persistent. One advantage of having a DOI instead of just an OSF record or OSF URL is that publishers, repositories, aggregators and other providers of research information use DOIs to identify research work precisely. data site and crossref not only maintain that PID, they also keep a catalog of all of the material that has DOIs. So all those publishers and repositories can use um can get access to that metadata and create more places your work gets disseminated. In addition to just creating DOIs for your OSF content, you can link to DOIs or other materials from your registrations. These show up as badges on the registrations. You can add DOIs for data, materials, analytic code, papers, and other supplements. Then when someone comes to your registration and learns about what you have planned for your project, they can immediately find out what your actual outputs and outcomes are. Creating those links is simple. From the registration overview page, you select any resource icon and those icons start out not colored in. Uh once you select the icon you want to um fill in, you see a window will open where you can add the resource by adding the DOI and and picking what type of content this is. And then once the resource is linked, it will show up with a colored icon. And you can add as many resources in these categories as you want. Multiple data sets, multiple papers, multiple uh types of code. You can also enrich enrich your preprints with PIDs for pre-registrations, supplemental materials, and published versions of papers. So remember this slide. Once you start adding PIDs like orchids and DOIs to your OSF content, you start to create your own graph of information about your research activities. [snorts] As an example, this is my orchid record and I have publications showing here that were added from the MLA international bibliography and Scopus. So, it's benefiting me to have used those pits to have registered my orchid um and set up that automatic uh update in my orchid record. So lastly, I've included this picture of my cat in every other OSF Essentials webinar I've done. So I felt like I had to include it before I ended today. Her name is Clementine, and she would like to know if you have any questions. So at this point, you can put questions, as I said, in the Q&A. I will give you a little bit of time to do that. And while we wait, I'll tell you about our upcoming OSF Essential sessions. So next week uh next week next month um we're going to talk about digital preservation and research data. Data is very easy to duplicate but it's actually very hard to preserve. So um we'll take a look at the basic needs of digital preservation and give you tips and strategies to ensure the longevity of your work. Then we'll talk about registration templates on the OSF in September and we'll talk about how to use the OSF API in October. We've also launched a new series of self-paced training opportunities that help researchers build practical uh skills across the research life cycle using open scholarship practices. Um the first three courses are available now. Fundamentals of open scholarship, pre-registration and registered reports and data management and sharing. And we'll have more courses coming throughout 2026. You can also get badges for the completion of this uh of these modules. So, you can find out more at cos.io/training/selfpaste. All right. So, I'm going to go back to our questions and if I can find it. Do we have any? [clears throat] So, the question in the chat is if you have an institutional membership with OSF and then end it. If someone while you had the membership had the roar included with their OSF account, will it remain there for metadata purposes? So uh no it won't show in OSF but the data site record won't be updated right away. So it will remain in data site for a while. Um but maintaining the the roar is part of the in in the record as part of the the membership um service. I see we have um we don't have any other questions at this time. I'll give folks another minute if they'd like to add anything. Just check if there was anything in the chat. Doesn't look like it. Okay. Well, we can go ahead and wrap up for today. This was a quick one, but um hopefully it was enough to show you how PIDs can help you um make the most of your research. So, like I said, you'll receive an email um either in the in the next few days, probably early next week, that will have a link to this uh recording as well as the slides and the links that were shared. So, thank you everybody. Have a great rest of your day. Bye.