Skip to content

Improving Your S-index

This page covers practical ways to grow your S-index. Note that these are intended to encourage good data management and sharing practices, not ways to inflate your S-index, since gaming can be easily detected as explained in Understanding Your S-index.

Choose Repositories That Expose Rich Metadata

The FAIR score is calculated using F-UJI, which relies on the dataset metadata exposed by a repository. So both the repository you choose and how much metadata you provide through it matter.

To improve your S-index, share your datasets on repositories that expose metadata properly, and fill in as much metadata as the repository allows when depositing your data.

To check how FAIR-enabling a repository is, you can look up the FAIR scores of some of its datasets on Scholar Data or select a few datasets from that repository and run them through the F-UJI web tool yourself.

Prefer a DOI-Issuing Repository

Citations and mentions are tracked more reliably for datasets that have a DOI. Much of the infrastructure that tracks reuse of scholarly output, including DataCite and OpenAlex, is built around DOIs, so a dataset without one is more likely to have real-world reuse go undetected. Prefer a DOI-issuing repository whenever you have the choice.

Cite Your Datasets Properly

When you use your own datasets in a manuscript, cite them as a proper reference rather than mentioning them only in text. This makes the citation trackable by the infrastructure that feeds your S-index, and sets a visible example that encourages others citing your work to do the same, which, over time, adds to your dataset's citation count.

Make Your Datasets Interoperable and Reusable

Datasets that are well organized, shared with standard file formats, and properly documented are easier for other researchers to reuse. Making your dataset as frictionless to reuse as possible can increase citations and mentions.

Share More Datasets

Sharing more datasets adds more Dataset Indices to your total S-index, but only if each one carries real impact. Sharing many low-impact datasets doesn't meaningfully move the needle, and is visible as such (see Design Rationale). Treat this as a natural result of consistently sharing well-prepared, reusable data, not a shortcut on its own.

Share Early

Citations and mentions are time-weighted, so the sooner a dataset is shared, the sooner it starts accumulating reuse, and the more that reuse counts toward your S-index over time. Waiting to share a dataset until a related manuscript is published, for example, delays this clock unnecessarily.

Documentation written with assistance from Claude by Anthropic.