API Reference

class dataplan.DataPlan(execplan, sources=None)

A plan to process data

aggregate(*args, **kwargs)

Perform a series of aggregations.

property execplan

The current plan to be executed.

filter(*args, **kwargs)

Apply a filter.

classmethod from_dataframe(df, **kwargs)

Construct a DataPlan from a pandas DataFrame.

Parameters:
  • df – pandas dataframe from which to construct table source

  • kwargs – optional keyword arguments to pa.table.from_pandas`

Returns:

the DataPlan.

classmethod from_dataset(dataset, columns=None, filter=None)

Construct a DataPlan from a pyarrow Dataset.

Parameters:
  • dataset – pyarrow Dataset to be declared as source

  • columns – Optional projection to apply

  • filter – Optional predicate to apply

Returns:

the DataPlan.

classmethod from_dict(mapping, **kwargs)

Construct a DataPlan from a dictionary.

Parameters:
  • mapping – Dictionary from which to construct Table source

  • kwargs – Optional keyword arguments to pa.Table.from_pydict`

Returns:

The DataPlan.

classmethod from_list(mapping, **kwargs)

Construct a DataPlan from a list.

Parameters:
  • mapping – List of dictionaries from which to construct Table source

  • kwargs – Optional keyword arguments to pa.Table.from_pylist`

Returns:

The DataPlan.

classmethod from_table(table)

Construct a DataPlan from a pyarrow Table.

Parameters:

table – pyarrow Table to be declared as source

Returns:

the DataPlan.

group_by(keys)

Group according to some keys.

join(other, *args, **kwargs)

Perform a join with another DataPlan.

join_asof(other, *args, **kwargs)

Perform an asof join with another DataPlan.

order_by(*args, **kwargs)

Apply an ordering.

project(*args, **kwargs)

Apply a projection (e.g., select columns, rename columns, construct new columns, etc.).

property sources

The sources over which the plan will operate.

to_array(use_threads=True)

Execute the plan and return a numpy array.

to_batches(use_threads=True)

Execute the plan and return a list of pyarrow RecordBatches.

to_dataframe(use_threads=True)

Execute the plan and return a pandas DataFrame.

to_dict(use_threads=True)

Execute the plan and return a dictionary of columns.

to_list(use_threads=True)

Execute the plan and return a list of rows.

to_reader(use_threads=True)

Execute the plan lazily as a reader.

to_recarray(use_threads=True)

Execute the plan and return a numpy record array.

to_table(use_threads=True)

Execute the plan and return a pyarrow Table.