The digital era has generated a huge amount of data on the identities (profiles) of people, organizations and other entities in a digital format, largely consisting of textual documents such as news articles, encyclopedias, personal websites, books, and social media. Identity has thus been transformed from a philosophical to a societal issue, one requiring robust computational tools to determine entity identity in text.
Computational systems developed to establish identity in text often struggle with long-tail cases. This book investigates how Natural Language Processing (NLP) techniques for establishing the identity of long-tail entities – which are all infrequent in communication, hardly represented in knowledge bases, and potentially very ambiguous – can be improved through the use of background knowledge. Topics covered include: distinguishing tail entities from head entities; assessing whether current evaluation datasets and metrics are representative for long-tail cases; improving evaluation of long-tail cases; accessing and enriching knowledge on long-tail entities in the Linked Open Data cloud; and investigating the added value of background knowledge (“profiling”) models for establishing the identity of NIL entities.
Providing novel insights into an under-explored and difficult NLP challenge, the book will be of interest to all those working in the field of entity identification in text.
Paga facilmente con carta, Klarna, Apple Pay o Google Pay. Non sei soddisfatto? Hai sempre 14 giorni per il rimborso. Leggi di più nei nostri termini. Per qualsiasi domanda, scrivici a hello@memmo.org.
Memmo rende lo studio più facile, ovunque tu sia nel mondo. Qui trovi i tuoi libri di testo e strumenti di studio intelligenti, tutto in un unico posto: riassunti, quiz, podcast e flashcard. E poi c'è Ted, il tuo compagno di studio che risponde a ogni tua domanda. Oltre 50.000 studenti studiano già qui: è stato creato per aiutarti a imparare più velocemente e a stressarti meno.