timeseries_backend
timeseries_backend
¶
Time-series generation backend with chronological validation.
Classes:
| Name | Description |
|---|---|
ProgressSnapshot |
Snapshot configuration for saving partial generation results at progress milestones. |
RecordPromptState |
Mutable prefix/history state used to build record prompts. |
GroupState |
Mutable state for tracking a single group during parallel generation. |
GroupProcessingResult |
Result of processing a generation batch for a single group. |
TimeseriesBackend |
Time-series aware generator that enforces chronological constraints. |
ProgressSnapshot(label, threshold, path, saved=False)
dataclass
¶
Snapshot configuration for saving partial generation results at progress milestones.
Attributes:
| Name | Type | Description |
|---|---|---|
label |
str
|
Human-readable label for the milestone (e.g. |
threshold |
int
|
Record or group count that triggers this snapshot. |
path |
Path
|
File path where the snapshot CSV will be written. |
saved |
bool
|
Whether this snapshot has already been written to disk. |
label
instance-attribute
¶
Human-readable label for the milestone (e.g. "50").
threshold
instance-attribute
¶
Record or group count that triggers this snapshot.
path
instance-attribute
¶
File path where the snapshot CSV will be written.
saved = field(default=False)
class-attribute
instance-attribute
¶
Whether this snapshot has already been written to disk.
RecordPromptState(prefix, history=list(), using_prefix=True)
dataclass
¶
Mutable prefix/history state used to build record prompts.
Methods:
| Name | Description |
|---|---|
add_history |
Accept records, switch from prefix to history, and clamp the window. |
Attributes:
| Name | Type | Description |
|---|---|---|
prefix |
str
|
Incomplete first record used until the first generated record is accepted. |
history |
list[ParsedRecord]
|
Exact accepted record text used as rolling prompt history. |
using_prefix |
bool
|
Whether generation is still completing the initial record prefix. |
prompt_segments |
str | Sequence[str]
|
Return the prefix or individually encoded history records. |
completion_prefix |
str
|
Return bytes prepended to completions during the prefix phase. |
history_text |
str
|
Return history as newline-terminated training-compatible JSONL. |
prefix
instance-attribute
¶
Incomplete first record used until the first generated record is accepted.
history = field(default_factory=list)
class-attribute
instance-attribute
¶
Exact accepted record text used as rolling prompt history.
using_prefix = True
class-attribute
instance-attribute
¶
Whether generation is still completing the initial record prefix.
prompt_segments
property
¶
Return the prefix or individually encoded history records.
completion_prefix
property
¶
Return bytes prepended to completions during the prefix phase.
history_text
property
¶
Return history as newline-terminated training-compatible JSONL.
add_history(records, *, max_records)
¶
Accept records, switch from prefix to history, and clamp the window.
Source code in src/nemo_safe_synthesizer/generation/timeseries_backend.py
GroupState(group_id, group_ordinal, prompt_state, expected_records=0, last_timestamp_seconds=None, low_valid_fraction_count=0, completed=False, failed=False, total_valid_records=0, total_invalid_records=0, no_progress_count=0)
dataclass
¶
Mutable state for tracking a single group during parallel generation.
Each group maintains its own sliding-window context, timestamp cursor, and retry counters so that multiple groups can be generated in parallel while tracking progress independently.
Attributes:
| Name | Type | Description |
|---|---|---|
group_id |
TimeSeriesGroupValue
|
Unique identifier for this group (e.g., device ID, customer ID). |
group_ordinal |
int
|
Stable one-based position in the saved group registry, used in diagnostics. |
prompt_state |
RecordPromptState
|
Prefix/history container for building prompts and parsing completions. |
expected_records |
int
|
Target record count, calculated from |
last_timestamp_seconds |
int | None
|
Timestamp (in seconds) of the most recently generated record, used for chronological validation. |
low_valid_fraction_count |
int
|
Consecutive batches with high invalid fraction. Triggers group failure after |
completed |
bool
|
Whether this group has reached the stop timestamp. |
failed |
bool
|
Whether this group failed (e.g., too many retries without progress). |
total_valid_records |
int
|
Cumulative count of valid records generated for this group. |
total_invalid_records |
int
|
Cumulative count of invalid records generated for this group. |
no_progress_count |
int
|
Consecutive batches that did not advance the accepted timestamp. |
group_id
instance-attribute
¶
Unique identifier for this group (e.g., device ID, customer ID).
group_ordinal
instance-attribute
¶
Stable one-based position in the saved group registry, used in diagnostics.
prompt_state
instance-attribute
¶
Prefix/history container for building prompts and parsing completions.
expected_records = 0
class-attribute
instance-attribute
¶
Target record count, calculated from (stop_timestamp - start_timestamp) / interval_seconds.
last_timestamp_seconds = None
class-attribute
instance-attribute
¶
Timestamp (in seconds) of the most recently generated record, used for chronological validation.
low_valid_fraction_count = 0
class-attribute
instance-attribute
¶
Consecutive batches with high invalid fraction. Triggers group failure after patience is exceeded.
completed = False
class-attribute
instance-attribute
¶
Whether this group has reached the stop timestamp.
failed = False
class-attribute
instance-attribute
¶
Whether this group failed (e.g., too many retries without progress).
total_valid_records = 0
class-attribute
instance-attribute
¶
Cumulative count of valid records generated for this group.
total_invalid_records = 0
class-attribute
instance-attribute
¶
Cumulative count of invalid records generated for this group.
no_progress_count = 0
class-attribute
instance-attribute
¶
Consecutive batches that did not advance the accepted timestamp.
GroupProcessingResult
¶
Bases: Enum
Result of processing a generation batch for a single group.
Used by _process_group_result to signal whether a group should
remain active, be marked complete, or be removed due to failure.
Attributes:
| Name | Type | Description |
|---|---|---|
IN_PROGRESS |
Group continues; batch should be added to the accumulator. |
|
COMPLETED |
Group reached the stop timestamp; remove from active processing. |
|
FAILED |
Group failed after exhausting a consecutive retry condition. |
IN_PROGRESS = auto()
class-attribute
instance-attribute
¶
Group continues; batch should be added to the accumulator.
COMPLETED = auto()
class-attribute
instance-attribute
¶
Group reached the stop timestamp; remove from active processing.
FAILED = auto()
class-attribute
instance-attribute
¶
Group failed after exhausting a consecutive retry condition.
TimeseriesBackend(config, model_metadata, **kwargs)
¶
Bases: VllmBackend
Time-series aware generator that enforces chronological constraints.
This backend extends VllmBackend to generate synthetic time-series data with strict chronological ordering. It uses a sliding window approach where recently generated records are used as history for subsequent generation, ensuring temporal continuity.
Key Concepts
- Time-Range Based Generation: The number of records generated is
determined by the configured time range and interval, not by a target
count. Specifically: (stop_timestamp - start_timestamp) / interval_seconds.
The
config.generation.num_recordsparameter is used only for progress tracking, not to limit output. - Sliding Window: The backend maintains a window of recent records
(controlled by
_history_window_size) that are included in each prompt to provide context for the LLM, ensuring generated records follow the established patterns and timestamps. - Parallel Group Generation: Multiple time-series groups (e.g., different
devices, customers) are processed in parallel batches for efficiency.
Even single-sequence data uses this path (treated as 1 group via a
pseudo-group column added during preprocessing). Groups are the same as
those seen during training (from
model_metadata.timeseries_group_values). - Chronological Validation: Each generated record must continue from the previous timestamp at the expected interval. Out-of-order records are marked invalid.
Generation Flow (parallel group mode): 1. Initialize GroupState for each group with a partial first record 2. While groups remain pending or active: a. Fill active slots with pending groups (up to max_groups_per_batch) b. Build prompts for all active groups using their prefix or history c. Generate completions for all prompts in a single LLM batch call d. Process LLM outputs into per-group Batch objects e. For each group: - Validate chronological order against group's last timestamp - Retain the response with the most valid records (discard others) - Apply data actions and discard records rejected by post-processing - Update group state from accepted records (history, last_timestamp) - Check if an accepted record reached the stop timestamp - Track invalid output and timestamp progress; fail after consecutive retries f. Remove completed/failed groups from active list g. Save progress snapshots if thresholds are met h. Log per-group progress summary
Stopping Conditions
Generation stops when all groups finish (either completed or failed). Individual groups and the overall generation can stop for different reasons:
Per-Group Stopping:
- Completion (success): A group completes when any generated record
has a timestamp >= _stop_timestamp_value. The group is marked as
completed and removed from active processing.
- Failure (low valid fraction or no progress): A group fails after
config.generation.patience consecutive batches where either the
invalid record fraction remains above the configured threshold or
no accepted timestamp advances. Failed groups are not retried;
records accepted in earlier batches remain in the partial output.
Global Stopping:
- Natural completion: Generation ends when both the pending groups
queue and active groups list are empty (all groups processed).
- No records: If GenerationBatches detects too many consecutive
batches with no valid records globally, it signals STOP_NO_RECORDS.
- Target reached: If the target number of records is reached,
GenerationBatches signals STOP_METRIC_REACHED.
When global stopping occurs before all groups complete, all_groups_succeeded
returns False, and the final generation status reflects partial completion.
Attributes:
| Name | Type | Description |
|---|---|---|
_samples_per_prompt |
int
|
Number of completion samples to generate per prompt. Multiple samples increase chances of getting valid records. Default: 5. |
_max_prompts_per_batch |
int
|
Maximum number of prompts to include in a single LLM generation call. Controls parallelism. Default: 100. |
_history_window_size |
int
|
Number of recent records to include in the sliding prompt history. Default: 3. |
_time_column |
str
|
Name of the timestamp column in the data. |
_time_format |
str
|
Format string for parsing timestamps (strptime format), or "elapsed_seconds" for numeric elapsed time. |
_is_elapsed_time |
bool
|
True if timestamps are numeric elapsed seconds. |
_start_timestamp_value |
Starting timestamp for generation range. |
|
_stop_timestamp_value |
Ending timestamp for generation range. Generation stops when a record reaches or exceeds this timestamp. |
|
_timestamp_interval_seconds |
int | None
|
Expected interval between consecutive timestamps. Used for chronological validation. |
_group_column |
str
|
Column name used to group time-series data. |
_group_ordinals |
dict[TimeSeriesGroupValue, int]
|
Mapping of group IDs to stable, non-sensitive positions used in diagnostics. |
_group_prefixes |
dict[TimeSeriesGroupValue, str]
|
Mapping of group IDs to incomplete first records used to start generation. |
_groups |
list[TimeSeriesGroupValue]
|
Typed group IDs to generate. |
Methods:
| Name | Description |
|---|---|
generate |
Generate time-series tabular data using Nemo Safe Synthesizer. |
Source code in src/nemo_safe_synthesizer/generation/timeseries_backend.py
generate(data_actions_fn=None)
¶
Generate time-series tabular data using Nemo Safe Synthesizer.
All time series are processed as groups (single-sequence is treated as 1 group via pseudo-group column added during preprocessing).
Note
Generation is time-range based, not count-based. The number of records generated is determined by (stop_timestamp - start_timestamp) / interval_seconds for each group. The config.generation.num_records parameter is used for progress tracking but does not limit output. Groups are the same as those seen during training (from model_metadata.timeseries_group_values).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data_actions_fn
|
DataActionsFn | None
|
Optional function that takes a DataFrame and returns a modified DataFrame. |
None
|
Returns:
| Type | Description |
|---|---|
GenerateJobResults
|
Generation results object, which includes a DataFrame of generated records. |
Source code in src/nemo_safe_synthesizer/generation/timeseries_backend.py
1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 | |