Lemmatization errors when text contains contracted forms of 'be' #674

gppatt · 2016-12-09T22:49:09Z

I've noticed some inconsistent behavior here:

nlp = spacy_nlp(u"I'm hungry. You're hungry. He's hungry. It's hungry. We're hungry. They're hungry.")
for tok in nlp:
print tok, tok.lemma_

Gives output:

I i
'm be
hungry hungry
. .
You you
're 're
hungry hungry
. .
He he
's '
hungry hungry
. .
It it
's '
hungry hungry
. .
We we
're 're
hungry hungry
. .
They they
're 're
hungry hungry
. .

A related error is for "won't" (and for the much rarer "shan't"):

nlp = spacy_nlp(u"They won't move.")
for tok in nlp:
print tok, tok.lemma_

They they
wo wo
n't not
move move
. .

I think I once even saw a similar lemmatization error for "can't", but I am not able to recreate this error.

Your Environment

OSX 10.11.6
Spyder 3.0.0
spaCy 1.2.0

ines · 2016-12-10T01:39:12Z

Thanks for the report! Some of these should probably be handled in the morphological analyser (like "He's", where the lemma is ambiguous). But the others are definitely cases for the TOKENIZER_EXCEPTIONS.

I'm currently in the process of reorganising the language data (see organize-language-data branch). I'll add the missing exceptions, so this should all be fixed in the v2.0 release.

lock · 2018-05-09T05:38:59Z

This thread has been automatically locked since there has not been any recent activity after it was closed. Please open a new issue for related bugs.

ines added performance 🌙 nightly Discussion and contributions related to nightly builds labels Dec 10, 2016

ines added this to the Reorganise language data milestone Dec 10, 2016

ines closed this as completed in a223221 Dec 18, 2016

ines removed the 🌙 nightly Discussion and contributions related to nightly builds label Dec 18, 2016

lock bot locked as resolved and limited conversation to collaborators May 9, 2018

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Lemmatization errors when text contains contracted forms of 'be' #674

Lemmatization errors when text contains contracted forms of 'be' #674

gppatt commented Dec 9, 2016

ines commented Dec 10, 2016 •

edited

Loading

lock bot commented May 9, 2018

Lemmatization errors when text contains contracted forms of 'be' #674

Lemmatization errors when text contains contracted forms of 'be' #674

Comments

gppatt commented Dec 9, 2016

Your Environment

ines commented Dec 10, 2016 • edited Loading

lock bot commented May 9, 2018

ines commented Dec 10, 2016 •

edited

Loading