Python interface to the Anserini IR toolkit built on Lucene
Project description
Pyserini provides a simple Python interface to the Anserini IR toolkit via pyjnius.
Installation
Install via PyPI
pip install pyserini
Usage
Here's a sample pre-built index on TREC Disks 4 & 5 to play with (used in the TREC 2004 Robust Track):
wget https://git.uwaterloo.ca/jimmylin/anserini-indexes/raw/master/index-robust04-20191213.tar.gz
tar xvfz index-robust04-20191213.tar.gz
Use the SimpleSearcher
for searching:
from pyserini.search import pysearch
searcher = pysearch.SimpleSearcher('index-robust04-20191213/')
hits = searcher.search('hubble space telescope')
# Print the first 10 hits:
for i in range(0, 10):
print(f'{i+1} {hits[i].docid} {hits[i].score}')
# Grab the actual text:
hits[0].content
For additional information, please refer to the Pyserini repository.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
pyserini-0.8.1.0.tar.gz
(57.7 MB
view hashes)
Built Distribution
pyserini-0.8.1.0-py3-none-any.whl
(57.7 MB
view hashes)
Close
Hashes for pyserini-0.8.1.0-py3-none-any.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | 8755c071b67fd815809e163304c0f0219286ae9b268893ff16b4d226775eec6e |
|
MD5 | 6a0efb6142486bdf6c0e8bdbe89e6ad9 |
|
BLAKE2b-256 | 72f445a9175211da4b7d4aed77af707eab938fc068d6124fc6673da748bc74bd |