Type of site | Search engine |
|---|---|
| Created by | Allen Institute for Artificial Intelligence |
| URL | semanticscholar |
| Launched | November 2, 2015; 9 years ago (2015-11-02)[1] |
Semantic Scholar is a research tool for scientific literature. It is developed at theAllen Institute for AI and was publicly released in November 2015.[2] Semantic Scholar uses modern techniques innatural language processing to support the research process, for example by providing automatically generated summaries of scholarly papers.[3] The Semantic Scholar team is actively researching the use of artificial intelligence innatural language processing,machine learning,human–computer interaction, andinformation retrieval.[4]
Semantic Scholar began as a database for the topics ofcomputer science,geoscience, andneuroscience.[5] In 2017, the system began includingbiomedical literature in its corpus.[5] As of September 2022[update], it includes over 200 million publications from all fields of science.[6]
Semantic Scholar provides a one-sentence summary ofscientific literature. One of its aims was to address the challenge of reading numerous titles and lengthy abstracts on mobile devices.[7] It also seeks to ensure that the three million scientific papers published yearly reach readers, since it is estimated that only half of this literature is ever read.[8]
Artificial intelligence is used to capture the essence of a paper, generating it through an "abstractive" technique.[3] The project uses a combination ofmachine learning,natural language processing, andmachine vision to add a layer ofsemantic analysis to the traditional methods ofcitation analysis, and to extract relevant figures,tables, entities, and venues from papers.[9][10]
Another key AI-powered feature is Research Feeds, an adaptive research recommender that uses AI to quickly learn what papers users care about reading and recommends the latest research to help scholars stay up to date. It uses a state-of-the-art paper embedding model trained using contrastive learning to find papers similar to those in each Library folder.[11]
Semantic Scholar also offers Semantic Reader, an augmented reader with the potential to revolutionize scientific reading by making it more accessible and richly contextual.[12] Semantic Reader provides in-line citation cards that allow users to see citations withTLDR (short for Too Long, Didn't Read) automatically generated short summaries as they read and skimming highlights that capture key points of a paper so users can digest faster.
In contrast withGoogle Scholar andPubMed, Semantic Scholar is designed to highlight the most important and influential elements of a paper.[13] The AI technology is designed to identify hidden connections and links between research topics.[14] Like the previously cited search engines, Semantic Scholar also exploits graph structures, which include theMicrosoft Academic Knowledge Graph, Springer Nature'sSciGraph, and the Semantic Scholar Corpus (originally a 45 million papers corpus in computer science, neuroscience and biomedicine).[15][16]
Each paper hosted by Semantic Scholar is assigned a uniqueidentifier called the Semantic Scholar Corpus ID (abbreviated S2CID). The following entry is an example:
Liu, Ying; Gayle, Albert A; Wilder-Smith, Annelies; Rocklöv, Joacim (March 2020). "The reproductive number of COVID-19 is higher compared to SARS coronavirus".Journal of Travel Medicine.27 (2).doi:10.1093/jtm/taaa021.PMID 32052846.S2CID 211099356.
Semantic Scholar is free to use and unlike similar search engines (i.e.Google Scholar) does not search for material that is behind apaywall.[5][better source needed]
One study compared the index scope of Semantic Scholar to Google Scholar, and found that for the papers cited by secondary studies in computer science, the two indices had comparable coverage, each only missing a handful of the papers.[17]
As of January 2018, following a 2017 project that added biomedical papers and topic summaries, the Semantic Scholar corpus included more than 40 million papers fromcomputer science andbiomedicine.[18] In March 2018, Doug Raymond, who developedmachine learning initiatives for theAmazon Alexa platform, was hired to lead the Semantic Scholar project.[19] As of August 2019[update], the number of included papers metadata (not the actual PDFs) had grown to more than 173 million[20] after the addition of theMicrosoft Academic Graph records.[21] In 2020, a partnership between Semantic Scholar and theUniversity of Chicago Press Journals made all articles published under the University of Chicago Press available in the Semantic Scholar corpus.[22] At the end of 2020, Semantic Scholar had indexed 190 million papers.[23] In 2020, Semantic Scholar reached seven million users per month.[7]
...the publicly available corpus compiled by Semantic Scholar – a tool set up in 2015 by the Allen Institute for Artificial Intelligence in Seattle, Washington – amounting to around 200 million articles, including preprints.