NLP, before and after spaCy — textacy 0.13.0 documentation

textacy is a Python library for performing a variety of natural language processing (NLP)
tasks, built on the high-performance spaCy library. With the fundamentals — tokenization,
part-of-speech tagging, dependency parsing, etc. — delegated to another library,
textacy focuses primarily on the tasks that come before and follow after.

features

Access and extend spaCy’s core functionality for working with one or many documents
through convenient methods and custom extensions
Load prepared datasets with both text content and metadata, from Congressional speeches
to historical literature to Reddit comments
Clean, normalize, and explore raw text before processing it with spaCy
Extract structured information from processed documents, including n-grams, entities,
acronyms, keyterms, and SVO triples
Compare strings and sequences using a variety of similarity metrics
Tokenize and vectorize documents then train, interpret, and visualize topic models
Compute text readability and lexical diversity statistics, including Flesch-Kincaid
grade level, multilingual Flesch Reading Ease, and Type-Token Ratio

… and much more!

maintainer

Howdy, y’all. 👋

Source link

Latest articles

NLP, before and after spaCy — textacy 0.13.0 documentation

features

maintainer

Latest articles

ChatGPT gained one million new users in an hour today

China police deploy real-life Robocop as humanoid tech takes huge leap forward

Runway releases Gen-4 video model with focus on consistency

Leave a Comment Cancel reply

Featured articles

ChatGPT gained one million new users in an hour today

China police deploy real-life Robocop as humanoid tech takes huge leap forward

Runway releases Gen-4 video model with focus on consistency