timeseries_prompting
timeseries_prompting
¶
Prompt construction helpers for time-series generation.
Functions:
| Name | Description |
|---|---|
build_partial_record_prefix |
Build an incomplete first record matching the training serialization. |
build_record_history |
Build prompt history from exact model-emitted record text. |
build_training_compatible_prompt_token_ids |
Build the token prefix seen before record continuation during training. |
build_partial_record_prefix(*, columns, schema, group_column, group_id, timestamp_column, start_timestamp)
¶
Build an incomplete first record matching the training serialization.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
columns
|
Sequence[str]
|
Saved JSON schema columns in generation order. |
required |
schema
|
Mapping[str, object]
|
Saved JSON schema used to coerce known values. |
required |
group_column
|
str
|
Configured group column, including the pseudo-group value. |
required |
group_id
|
object
|
Group value for this generation stream. |
required |
timestamp_column
|
str
|
Configured timestamp column. |
required |
start_timestamp
|
str | int
|
First timestamp to generate. |
required |
Returns:
| Type | Description |
|---|---|
str
|
An incomplete JSON record ending with the opening quote of the next |
str
|
field name. Including the training |
str
|
prefix tokenization identical to the beginning of a complete training |
str
|
record. The record begins directly with |
str
|
the sequence BOS token immediately before the first JSON byte. |
Examples:
Given columns beginning with acct_id, txn_index, and
cardholder, a group ID of "ACCT-001", and a starting transaction
index of 1, the returned prefix is::
{"acct_id":"ACCT-001","txn_index":1,"
The model completes the next field name and the rest of the record.
Raises:
| Type | Description |
|---|---|
GenerationError
|
If the saved artifact cannot support a partial prefix. |
Source code in src/nemo_safe_synthesizer/generation/timeseries_prompting.py
build_record_history(records)
¶
Build prompt history from exact model-emitted record text.
Training inserts the sequence BOS token immediately before the first record, with no intervening whitespace. The caller adds that BOS token separately, so this helper returns only the exact newline-terminated record bytes.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
records
|
Sequence[ParsedRecord]
|
Accepted records in chronological order. |
required |
Returns:
| Type | Description |
|---|---|
str
|
A sequence of newline-terminated records, or an empty string when no |
str
|
records are provided. |
Source code in src/nemo_safe_synthesizer/generation/timeseries_prompting.py
build_training_compatible_prompt_token_ids(*, tokenizer, prompt_config, prompt, record_context)
¶
Build the token prefix seen before record continuation during training.
This mirrors Example construction: encode the schema prompt without
tokenizer-defined special tokens, apply the configured prompt BOS/EOS
tokens, add the sequence BOS token, then append the prefix or history
record context. Building IDs explicitly is necessary for tokenizers such as
SmolLM3's, which do not add <|im_start|> automatically at inference.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tokenizer
|
EncodeOnlyTokenizer
|
Tokenizer used by the generation engine. |
required |
prompt_config
|
LLMPromptConfig
|
Saved special-token settings. |
required |
prompt
|
str
|
Schema prompt text used during training. |
required |
record_context
|
str | Sequence[str]
|
Partial first record, or separately encoded history records. Encoding history records individually mirrors training, which tokenizes each newline-terminated JSON record before concatenation. |
required |
Returns:
| Type | Description |
|---|---|
list[int]
|
Prompt token IDs whose boundary exactly matches a training example. |