Skip to main content

Simple, Pythonic text processing. Sentiment analysis, POS tagging, noun phrase parsing, and more.

Project description

https://travis-ci.org/sloria/TextBlob.png

Simplified text processing for Python 2 and 3.

Requirements

  • Python >= 2.6 or >= 3.3

Usage

Simple.

Create a TextBlob

from text.blob import TextBlob

zen = """Beautiful is better than ugly.
Explicit is better than implicit.
Simple is better than complex.
Complex is better than complicated.
Flat is better than nested.
Sparse is better than dense.
Readability counts.
Special cases aren't special enough to break the rules.
Although practicality beats purity.
Errors should never pass silently.
Unless explicitly silenced.
In the face of ambiguity, refuse the temptation to guess.
There should be one-- and preferably only one --obvious way to do it.
Although that way may not be obvious at first unless you're Dutch.
Now is better than never.
Although never is often better than *right* now.
If the implementation is hard to explain, it's a bad idea.
If the implementation is easy to explain, it may be a good idea.
Namespaces are one honking great idea -- let's do more of those!
"""

blob = TextBlob(zen) # Create a new TextBlob

Part-of-speech tags and noun phrases…

...are just properties.

blob.pos_tags         # [('beautiful', 'JJ'), ('is', 'VBZ'), ('better', 'RBR'),
                      # ('than', 'IN'), ('ugly', 'RB'), ...]

blob.noun_phrases     # ['beautiful', 'explicit', 'simple', 'complex', 'flat',
                      # 'sparse', 'readability', 'special cases',
                      # 'practicality beats purity', 'errors', 'unless',
                      # 'obvious way','dutch', 'right now', 'bad idea',
                      # 'good idea', 'namespaces', 'great idea']

Sentiment analysis

The sentiment property returns a tuple of the form (polarity, subjectivity) where polarity ranges from -1.0 to 1.0 and subjectivity ranges from 0.0 to 1.0.

blob.sentiment        # (0.20, 0.58)

Tokenization

blob.words            # WordList(['Beautiful', 'is', 'better'...'more',
                      #           'of', 'those'])

blob.sentences        # [Sentence('Beautiful is better than ugly.'),
                      #  Sentence('Explicit is better than implicit.'),
                      #  ...]

Get word and noun phrase frequencies

blob.word_counts['special']   # 2 (not case-sensitive by default)
blob.words.count('special')   # Same thing
blob.words.count('special', case_sensitive=True)  # 1

blob.noun_phrases.count('great idea')  # 1

TextBlobs are like Python strings!

blob[0:19]            # TextBlob("Beautiful is better")
blob.upper()          # TextBlob("BEAUTIFUL IS BETTER THAN UGLY...")
blob.find("purity")   # 293

apple_blob = TextBlob('apples')
banana_blob = TextBlob('bananas')
apple_blob < banana_blob           # True
apple_blob + ' and ' + banana_blob # TextBlob('apples and bananas')
"{0} and {1}".format(apple_blob, banana_blob)  # 'apples and bananas'

Get start and end indices of sentences

Use sentence.start and sentence.end. This can be useful for sentence highlighting, for example.

for sentence in blob.sentences:
    print(sentence)  # Beautiful is better than ugly
    print("---- Starts at index {}, Ends at index {}"\
                .format(sentence.start, sentence.end))  # 0, 30

Get a JSON-serialized version of the blob

blob.json   # '[{"sentiment": [0.2166666666666667, ' '0.8333333333333334],
            # "stripped": "beautiful is better than ugly", '
            # '"noun_phrases": ["beautiful"], "raw": "Beautiful is better than ugly. ", '
            # '"end_index": 30, "start_index": 0}
            #  ...]'

Installation

If you have pip:

pip install textblob

Or (if you must):

easy_install textblob

IMPORTANT: TextBlob depends on some NLTK models to work. The easiest way to get these is to run the download_corpora.py script included with this distribution. You can get it here . Then run:

python download_corpora.py

Testing

Run

nosetests

to run all tests.

License

TextBlob is licenced under the MIT license. See the bundled LICENSE file for more details.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

textblob-0.3.3.tar.gz (1.3 MB view details)

Uploaded Source

File details

Details for the file textblob-0.3.3.tar.gz.

File metadata

  • Download URL: textblob-0.3.3.tar.gz
  • Upload date:
  • Size: 1.3 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No

File hashes

Hashes for textblob-0.3.3.tar.gz
Algorithm Hash digest
SHA256 d9d1be47d97314b0e8a8d4cdb27b80df6edd11241d3e111496340f8a2f1c4ea3
MD5 2d8f6c1950c2326a40f3555d15761071
BLAKE2b-256 db80dbb8ce34507f37d8a912eb2b3bf7b13a2441ba367eb581833ba39711d91a

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page