API Reference¶
- class dataplan.DataPlan(execplan, sources=None)¶
A plan to process data
- aggregate(*args, **kwargs)¶
Perform a series of aggregations.
- property execplan¶
The current plan to be executed.
- filter(*args, **kwargs)¶
Apply a filter.
- classmethod from_dataframe(df, **kwargs)¶
Construct a DataPlan from a pandas DataFrame.
- Parameters:
df – pandas dataframe from which to construct table source
kwargs – optional keyword arguments to pa.table.from_pandas`
- Returns:
the DataPlan.
- classmethod from_dataset(dataset, columns=None, filter=None)¶
Construct a DataPlan from a pyarrow Dataset.
- Parameters:
dataset – pyarrow Dataset to be declared as source
columns – Optional projection to apply
filter – Optional predicate to apply
- Returns:
the DataPlan.
- classmethod from_dict(mapping, **kwargs)¶
Construct a DataPlan from a dictionary.
- Parameters:
mapping – Dictionary from which to construct Table source
kwargs – Optional keyword arguments to pa.Table.from_pydict`
- Returns:
The DataPlan.
- classmethod from_list(mapping, **kwargs)¶
Construct a DataPlan from a list.
- Parameters:
mapping – List of dictionaries from which to construct Table source
kwargs – Optional keyword arguments to pa.Table.from_pylist`
- Returns:
The DataPlan.
- classmethod from_table(table)¶
Construct a DataPlan from a pyarrow Table.
- Parameters:
table – pyarrow Table to be declared as source
- Returns:
the DataPlan.
- group_by(keys)¶
Group according to some keys.
- join(other, *args, **kwargs)¶
Perform a join with another DataPlan.
- join_asof(other, *args, **kwargs)¶
Perform an asof join with another DataPlan.
- order_by(*args, **kwargs)¶
Apply an ordering.
- project(*args, **kwargs)¶
Apply a projection (e.g., select columns, rename columns, construct new columns, etc.).
- property sources¶
The sources over which the plan will operate.
- to_array(use_threads=True)¶
Execute the plan and return a numpy array.
- to_batches(use_threads=True)¶
Execute the plan and return a list of pyarrow RecordBatches.
- to_dataframe(use_threads=True)¶
Execute the plan and return a pandas DataFrame.
- to_dict(use_threads=True)¶
Execute the plan and return a dictionary of columns.
- to_list(use_threads=True)¶
Execute the plan and return a list of rows.
- to_reader(use_threads=True)¶
Execute the plan lazily as a reader.
- to_recarray(use_threads=True)¶
Execute the plan and return a numpy record array.
- to_table(use_threads=True)¶
Execute the plan and return a pyarrow Table.