ds_provider_mock_py_lib.dataset

File: __init__.py Region: ds_provider_mock_py_lib/dataset

Mock Dataset

This module implements a read-only synthetic dataset for e2e testing.

Example

>>> from uuid import uuid4
>>> from ds_provider_mock_py_lib.dataset import MockDataset, MockDatasetSettings
>>> from ds_provider_mock_py_lib.linked_service import MockLinkedService, MockLinkedServiceSettings
>>> linked_service = MockLinkedService(
...     id=uuid4(),
...     name="mock-ls",
...     version="1.0.0",
...     settings=MockLinkedServiceSettings(),
... )
>>> dataset = MockDataset(
...     id=uuid4(),
...     name="mock-ds",
...     version="1.0.0",
...     linked_service=linked_service,
...     settings=MockDatasetSettings(row_count=5),
... )
>>> linked_service.connect()
>>> dataset.read()
>>> data = dataset.output

Submodules

Classes

MockDataset

Read-only tabular dataset that emits deterministic synthetic rows.

MockColumn

One synthetic column in a mock dataset.

MockDatasetSettings

Settings that define mock read scope and injected failure behaviour.

Package Contents

class ds_provider_mock_py_lib.dataset.MockDataset[source]

Bases: ds_resource_plugin_py_lib.common.resource.dataset.TabularDataset[ds_provider_mock_py_lib.linked_service.mock.MockLinkedService, ds_provider_mock_py_lib.dataset.settings.MockDatasetSettings, ds_resource_plugin_py_lib.common.serde.serialize.PandasSerializer, ds_resource_plugin_py_lib.common.serde.deserialize.PandasDeserializer]

Read-only tabular dataset that emits deterministic synthetic rows.

linked_service: ds_provider_mock_py_lib.linked_service.mock.MockLinkedService
settings: ds_provider_mock_py_lib.dataset.settings.MockDatasetSettings
serializer: ds_resource_plugin_py_lib.common.serde.serialize.PandasSerializer | None
deserializer: ds_resource_plugin_py_lib.common.serde.deserialize.PandasDeserializer | None
__post_init__() None[source]

Ensure serializer and deserializer are always available.

property supports_checkpoint: bool

Whether this dataset supports incremental loads via self.checkpoint.

Returns:

Always True for the mock provider.

Return type:

bool

property type: ds_provider_mock_py_lib.enums.ResourceType

Get the type of the dataset.

Returns:

ResourceType

read() None[source]

Read the current mock batch into self.output.

An empty checkpoint performs a full load (batch 0). A populated incremental watermark reads the next change batch. A pagination cursor resumes an in-flight batch after failure.

Raises:
  • ReadError – If the mock backend fails or settings are invalid for read.

  • ConnectionError – If dataset-level injection is configured as a connection error.

create() NoReturn[source]

Create is not supported by this dataset.

update() NoReturn[source]

Update is not supported by this dataset.

upsert() NoReturn[source]

Upsert is not supported by this dataset.

delete() NoReturn[source]

Delete is not supported by this dataset.

purge() NoReturn[source]

Purge is not supported by this dataset.

rename() NoReturn[source]

Rename is not supported by this dataset.

list() NoReturn[source]

List is not supported by this dataset.

close() None[source]

Close the dataset and underlying linked service.

_unsupported(method: str) NoReturn[source]

Raise NotSupportedError for a read-only method.

Parameters:

method – Dataset method name.

Raises:

NotSupportedError – Always.

class ds_provider_mock_py_lib.dataset.MockColumn[source]

Bases: ds_common_serde_py_lib.Serializable

One synthetic column in a mock dataset.

name: str

Column name emitted in self.output.

kind: ds_provider_mock_py_lib.enums.ColumnKind

Value generator used for this column.

value: Any = None

Payload for constant and enum.

A scalar is emitted on every row when kind is constant. A non-empty list is sampled when kind is enum.

prefix: str = ''

Prefix used when kind is text.

low: int = 0

Inclusive lower bound for random numeric kinds.

high: int = 1000

Exclusive upper bound for random numeric kinds.

null_every: int | None = None

When set, every Nth id is None (stable across batches).

class ds_provider_mock_py_lib.dataset.MockDatasetSettings[source]

Bases: ds_resource_plugin_py_lib.common.resource.dataset.DatasetSettings

Settings that define mock read scope and injected failure behaviour.

columns: list[MockColumn]

Synthetic columns included in every emitted row.

row_count: int = 100

Rows in the full load (batch 0, all inserts).

incremental_insert_count: int = 0

New unique primary keys delivered in each incremental batch.

incremental_update_count: int = 0

Existing primary keys re-delivered with changed column values (row hash changes).

incremental_noop_count: int = 0

Existing primary keys re-delivered with an identical row (same hash).

page_size: int | None = None

Page size. None emits the whole batch in one page.

seed: int = 42

Determinism seed for id selection and random cell values.

page_delay_ms: int = 0

Artificial delay applied after each successful page.

op_column: str | None = None

Optional label column for intended row kind.

Gold merge uses primary key plus row hash, not this column. Leave unset so the label cannot change the hash of a noop row. Counts are always in operation.metadata["ops"].

modified_at_column: str = '_modified_at'

Column that stores the deterministic modified-at timestamp.

raise_on_page: int | None = None

1-based page number within the current batch that should fail. None disables.

raise_as: ds_provider_mock_py_lib.enums.RaiseAs

Contract exception class used when raise_on_page matches.

raise_error: ds_provider_mock_py_lib.models.MockError

Error spec used when raise_on_page matches.