Who’s Who in Your Digital Collection: Developing a Tool for Name Disambiguation and Identity Resolution

Carol Jean Godby, Patricia Hswe, Larry Jackson, Judith Klavans, Lev Ratinov, Dan Roth

Abstract


In the past twenty years, the problem space of automatically recognizing, extracting, classifying, and disambiguating named entities (e.g., the names of people, places, and organizations) from digitized text has received considerable attention in research produced by the library, computer science, and the computational linguistics communities. However, linking the output of these advances with the library community continues to be a challenge. This paper describes work being done by the University of Illinois, the Online Computer Library Center (OCLC), and the University of Maryland to develop, evaluate and link Named Entity Recognition (NER) and Entity Resolution with tools used for search and access. Name identification and extraction tools, particularly when integrated with a resolution into an authority file (e.g., WorldCat Identities, Wikipedia, etc.), can enhance reliable subject access for a document collection, improving document discoverability by end-users.

Full Text: PDF

Refbacks

  • There are currently no refbacks.


Humanities Division Logo The Division of the Humanities
1115 East 58th Street, Chicago IL 60637 / Tel 773.702.8512 / Fax 773.702.6305
The University of Chicago