petastorm

Petastorm is a library enabling the use of Parquet storage from Tensorflow, Pytorch, and other Python-based ML training frameworks.

These details have not been verified by PyPI

Project links

Homepage

Project description

Petastorm

Petastorm is an open source data access library developed at Uber ATG. This library enables single machine or distributed training and evaluation of deep learning models directly from datasets in Apache Parquet format. Petastorm supports popular Python-based machine learning (ML) frameworks such as Tensorflow, PyTorch, and PySpark. It can also be used from pure Python code.

Documentation web site: https://petastorm.readthedocs.io

Installation

pip install petastorm

There are several extra dependencies that are defined by the petastorm package that are not installed automatically. The extras are: tf, tf_gpu, torch, opencv, docs, test.

For example to trigger installation of GPU version of tensorflow and opencv, use the following pip command:

pip install petastorm[opencv,tf_gpu]

Generating a dataset

A dataset created using Petastorm is stored in Apache Parquet format. On top of a Parquet schema, petastorm also stores higher-level schema information that makes multidimensional arrays into a native part of a petastorm dataset.

Petastorm supports extensible data codecs. These enable a user to use one of the standard data compressions (jpeg, png) or implement her own.

Generating a dataset is done using PySpark. PySpark natively supports Parquet format, making it easy to run on a single machine or on a Spark compute cluster. Here is a minimalistic example writing out a table with some random data.

HelloWorldSchema = Unischema('HelloWorldSchema', [
   UnischemaField('id', np.int32, (), ScalarCodec(IntegerType()), False),
   UnischemaField('image1', np.uint8, (128, 256, 3), CompressedImageCodec('png'), False),
   UnischemaField('other_data', np.uint8, (None, 128, 30, None), NdarrayCodec(), False),
])


def row_generator(x):
   """Returns a single entry in the generated dataset. Return a bunch of random values as an example."""
   return {'id': x,
           'image1': np.random.randint(0, 255, dtype=np.uint8, size=(128, 256, 3)),
           'other_data': np.random.randint(0, 255, dtype=np.uint8, size=(4, 128, 30, 3))}


def generate_hello_world_dataset(output_url='file:///tmp/hello_world_dataset'):
   rows_count = 10
   rowgroup_size_mb = 256

   spark = SparkSession.builder.config('spark.driver.memory', '2g').master('local[2]').getOrCreate()
   sc = spark.sparkContext

   # Wrap dataset materialization portion. Will take care of setting up spark environment variables as
   # well as save petastorm specific metadata
   with materialize_dataset(spark, output_url, HelloWorldSchema, rowgroup_size_mb):

       rows_rdd = sc.parallelize(range(rows_count))\
           .map(row_generator)\
           .map(lambda x: dict_to_spark_row(HelloWorldSchema, x))

       spark.createDataFrame(rows_rdd, HelloWorldSchema.as_spark_schema()) \
           .coalesce(10) \
           .write \
           .mode('overwrite') \
           .parquet(output_url)

HelloWorldSchema is an instance of a Unischema object. Unischema is capable of rendering types of its fields into different framework specific formats, such as: Spark StructType, Tensorflow tf.DType and numpy numpy.dtype.
To define a dataset field, you need to specify a type, shape, a codec instance and whether the field is nullable for each field of the Unischema.
We use PySpark for writing output Parquet files. In this example, we launch PySpark on a local box (.master('local[2]')). Of course for a larger scale dataset generation we would need a real compute cluster.
We wrap spark dataset generation code with the materialize_dataset context manager. The context manager is responsible for configuring row group size at the beginning and write out petastorm specific metadata at the end.
The row generating code is expected to return a Python dictionary indexed by a field name. We use row_generator function for that.
dict_to_spark_row converts the dictionary into a pyspark.Row object while ensuring schema HelloWorldSchema compliance (shape, type and is-nullable condition are tested).
Once we have a pyspark.DataFrame we write it out to a parquet storage. The parquet schema is automatically derived from HelloWorldSchema.

Plain Python API

The petastorm.reader.Reader class is the main entry point for user code that accesses the data from an ML framework such as Tensorflow or Pytorch. The reader has multiple features such as:

Selective column readout
Multiple parallelism strategies: thread, process, single-threaded (for debug)
N-grams readout support
Row filtering (row predicates)
Shuffling
Partitioning for multi-GPU training
Local caching

Reading a dataset is simple using the petastorm.reader.Reader class which can be created using the petastorm.make_reader factory method:

from petastorm import make_reader

 with make_reader('hdfs://myhadoop/some_dataset') as reader:
    for row in reader:
        print(row)

hdfs://... and file://... are supported URL protocols.

Once a Reader is instantiated, you can use it as an iterator.

Tensorflow API

To hookup the reader into a tensorflow graph, you can use the tf_tensors function:

with make_reader('file:///some/localpath/a_dataset') as reader:
   row_tensors = tf_tensors(reader)
   with tf.Session() as session:
       for _ in range(3):
           print(session.run(row_tensors))

Alternatively, you can use new tf.data.Dataset API;

with make_reader('file:///some/localpath/a_dataset') as reader:
    dataset = make_petastorm_dataset(reader)
    iterator = dataset.make_one_shot_iterator()
    tensor = iterator.get_next()
    with tf.Session() as sess:
        sample = sess.run(tensor)
        print(sample.id)

Pytorch API

As illustrated in pytorch_example.py, reading a petastorm dataset from pytorch can be done via the adapter class petastorm.pytorch.DataLoader, which allows custom pytorch collating function and transforms to be supplied.

Be sure you have torch and torchvision installed:

pip install torchvision

The minimalist example below assumes the definition of a Net class and train and test functions, included in pytorch_example:

import torch
from petastorm.pytorch import DataLoader

torch.manual_seed(1)
device = torch.device('cpu')
model = Net().to(device)
optimizer = torch.optim.SGD(model.parameters(), lr=0.01, momentum=0.5)

def _transform_row(mnist_row):
    transform = transforms.Compose([
        transforms.ToTensor(),
        transforms.Normalize((0.1307,), (0.3081,))
    ])
    return (transform(mnist_row['image']), mnist_row['digit'])


transform = TransformSpec(_transform_row, removed_fields=['idx'])

with DataLoader(make_reader('file:///localpath/mnist/train', num_epochs=10,
                            transform_spec=transform), batch_size=64) as train_loader:
    train(model, device, train_loader, 10, optimizer, 1)
with DataLoader(make_reader('file:///localpath/mnist/test', num_epochs=10,
                            transform_spec=transform), batch_size=1000) as test_loader:
    test(model, device, test_loader)

PySpark and SQL

Using the Parquet data format, which is natively supported by Spark, makes it possible to use a wide range of Spark tools to analyze and manipulate the dataset. The example below shows how to read a Petastorm dataset as a Spark RDD object:

# Create a dataframe object from a parquet file
dataframe = spark.read.parquet(dataset_url)

# Show a schema
dataframe.printSchema()

# Count all
dataframe.count()

# Show a single column
dataframe.select('id').show()

SQL can be used to query a Petastorm dataset:

spark.sql(
   'SELECT count(id) '
   'from parquet.`file:///tmp/hello_world_dataset`').collect()

You can find a full code sample here: pyspark_hello_world.py,

Non Petastorm Parquet Stores

Petastorm can also be used to read data directly from Apache Parquet stores. To achieve that, use make_batch_reader (and not make_reader). The following table summarizes the differences make_batch_reader and make_reader functions.

make_reader	make_batch_reader
Only Petastorm datasets (created using materializes_dataset)	Any Parquet store (some native Parquet column types are not supported yet.
The reader returns one record at a time.	The reader returns batches of records. The size of the batch is not fixed and defined by Parquet row-group size.
Predicates passed to make_reader are evaluated per single row.	Predicates passed to make_batch_reader are evaluated per batch.

Troubleshooting

See the Troubleshooting page and please submit a ticket if you can’t find an answer.

Publications

Gruener, R., Cheng, O., and Litvin, Y. (2018) Introducing Petastorm: Uber ATG’s Data Access Library for Deep Learning. URL: https://eng.uber.com/petastorm/

How to Contribute

We prefer to receive contributions in the form of GitHub pull requests. Please send pull requests against the github.com/uber/petastorm repository.

If you are looking for some ideas on what to contribute, check out github issues and comment on the issue.
If you have an idea for an improvement, or you’d like to report a bug but don’t have time to fix it please a create a github issue.

To contribute a patch:

Break your work into small, single-purpose patches if possible. It’s much harder to merge in a large change with a lot of disjoint features.
Submit the patch as a GitHub pull request against the master branch. For a tutorial, see the GitHub guides on forking a repo and sending a pull request.
Include a detailed describtion of the proposed change in the pull request.
Make sure that your code passes the unit tests. You can find instructions how to run the unit tests here.
Add new unit tests for your code.

Thank you in advance for your contributions!

See the Development for development related information.

Project details

These details have not been verified by PyPI

Project links

Homepage

Release history Release notifications | RSS feed

0.13.1

Jan 2, 2026

0.13.0rc0 pre-release

Aug 11, 2025

0.12.2rc0 pre-release

Feb 3, 2023

0.12.1

Dec 16, 2022

0.12.1rc1 pre-release

Dec 16, 2022

0.12.0

Aug 25, 2022

0.12.0rc3 pre-release

Aug 24, 2022

0.12.0rc2 pre-release

Aug 24, 2022

0.12.0rc1 pre-release

Aug 24, 2022

0.11.5

Jul 28, 2022

0.11.5rc0 pre-release

Jul 27, 2022

0.11.4

Feb 19, 2022

0.11.4rc0 pre-release

Feb 15, 2022

0.11.3

Sep 4, 2021

0.11.3rc0 pre-release

Sep 3, 2021

0.11.2

Aug 3, 2021

0.11.2rc1 pre-release

Jul 30, 2021

0.11.1

Jun 2, 2021

0.11.1rc0 pre-release

May 30, 2021

0.11.0

May 14, 2021

0.11.0rc6 pre-release

May 13, 2021

0.11.0rc5 pre-release

May 13, 2021

0.10.0

Apr 28, 2021

0.10.0rc4 pre-release

Apr 28, 2021

0.9.8

Dec 15, 2020

0.9.8rc0 pre-release

Dec 12, 2020

0.9.7

Oct 28, 2020

0.9.7rc0 pre-release

Oct 27, 2020

0.9.6

Sep 26, 2020

0.9.6rc0 pre-release

Sep 24, 2020

0.9.5

Aug 27, 2020

0.9.5rc0 pre-release

Aug 25, 2020

0.9.4

Jul 28, 2020

0.9.4rc0 pre-release

Jul 28, 2020

0.9.3

Jul 24, 2020

0.9.3rc1 pre-release

Jul 24, 2020

0.9.3rc0 pre-release

Jul 23, 2020

0.9.2

Jun 4, 2020

0.9.2rc0 pre-release

Jun 3, 2020

0.9.1

May 13, 2020

0.9.1rc0 pre-release

May 12, 2020

0.9.0

Apr 24, 2020

0.9.0rc0 pre-release

Apr 22, 2020

0.8.2

Feb 3, 2020

0.8.2rc0 pre-release

Feb 1, 2020

0.8.1

Jan 31, 2020

0.8.1rc1 pre-release

Jan 29, 2020

0.8.1rc0 pre-release

Jan 25, 2020

0.8.0

Dec 5, 2019

0.8.0rc0 pre-release

Dec 4, 2019

0.7.7

Oct 23, 2019

0.7.7rc9 pre-release

Oct 22, 2019

0.7.7rc8 pre-release

Oct 19, 2019

0.7.7rc7 pre-release

Oct 19, 2019

0.7.7rc6 pre-release

Oct 16, 2019

0.7.7rc5 pre-release

Oct 9, 2019

0.7.7rc4 pre-release

Oct 9, 2019

0.7.7rc3 pre-release

Oct 9, 2019

0.7.7rc2 pre-release

Oct 8, 2019

0.7.7rc0 pre-release

Sep 4, 2019

0.7.6

Aug 29, 2019

0.7.6rc0 pre-release

Aug 27, 2019

0.7.5

Jun 10, 2019

0.7.5rc0 pre-release

Jun 8, 2019

0.7.4

May 15, 2019

This version

0.7.4rc4 pre-release

May 14, 2019

0.7.4rc3 pre-release

May 14, 2019

0.7.3

May 1, 2019

0.7.3rc0 pre-release

May 1, 2019

0.7.2

Apr 25, 2019

0.7.2rc1 pre-release

Apr 24, 2019

0.7.1

Apr 3, 2019

0.7.1rc0 pre-release

Apr 2, 2019

0.7.0

Mar 25, 2019

0.7.0rc1 pre-release

Mar 22, 2019

0.7.0rc0 pre-release

Mar 22, 2019

0.6.0

Feb 23, 2019

0.6.0rc0 pre-release

Feb 21, 2019

0.5.1

Dec 27, 2018

0.5.0

Dec 12, 2018

0.5.0rc1 pre-release

Dec 11, 2018

0.5.0rc0 pre-release

Dec 10, 2018

0.4.3

Oct 11, 2018

0.4.3rc0 pre-release

Oct 10, 2018

0.4.2

Sep 26, 2018

0.4.2rc0 pre-release

Sep 25, 2018

0.4.1

Sep 15, 2018

0.3.1

Aug 30, 2018

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

petastorm-0.7.4rc4.tar.gz (158.5 kB view details)

Uploaded May 14, 2019 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

petastorm-0.7.4rc4-py2.py3-none-any.whl (255.5 kB view details)

Uploaded May 14, 2019 Python 2Python 3

File details

Details for the file petastorm-0.7.4rc4.tar.gz.

File metadata

Download URL: petastorm-0.7.4rc4.tar.gz
Upload date: May 14, 2019
Size: 158.5 kB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/1.13.0 pkginfo/1.5.0.1 requests/2.21.0 setuptools/41.0.1 requests-toolbelt/0.9.1 tqdm/4.32.1 CPython/3.6.3

File hashes

Hashes for petastorm-0.7.4rc4.tar.gz
Algorithm	Hash digest
SHA256	`d4c281c7b9343e66808622ccd44c5bbd1b87ac46208ba323f92ad9fbc4a580d2`
MD5	`97bb3ae88f340b0f6a6876a6d55c3bae`
BLAKE2b-256	`4d03fe4ce939274aaf1e52e7574e8524befa24706dc2f9e0f6822bb9b1102a7c`

See more details on using hashes here.

File details

Details for the file petastorm-0.7.4rc4-py2.py3-none-any.whl.

File metadata

Download URL: petastorm-0.7.4rc4-py2.py3-none-any.whl
Upload date: May 14, 2019
Size: 255.5 kB
Tags: Python 2, Python 3
Uploaded using Trusted Publishing? No
Uploaded via: twine/1.13.0 pkginfo/1.5.0.1 requests/2.21.0 setuptools/41.0.1 requests-toolbelt/0.9.1 tqdm/4.32.1 CPython/3.6.3

File hashes

Hashes for petastorm-0.7.4rc4-py2.py3-none-any.whl
Algorithm	Hash digest
SHA256	`c3cbabba4ba8086d4bd698adefaae01541773a7f622fa6390d267c0bdf43aa42`
MD5	`1b7ea777670512a243733dd915f36a4c`
BLAKE2b-256	`389e6c2c4998a376a41a97452ffbd6476d3559e78d616de1264e3919f2af66a7`

See more details on using hashes here.

petastorm 0.7.4rc4

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

Petastorm

Installation

Generating a dataset

Plain Python API

Tensorflow API

Pytorch API

PySpark and SQL

Non Petastorm Parquet Stores

Troubleshooting

Publications

How to Contribute

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes