Skip to content

timeseries_prompting

timeseries_prompting

Prompt construction helpers for time-series generation.

Functions:

Name Description
build_partial_record_prefix

Build an incomplete first record matching the training serialization.

build_record_history

Build prompt history from exact model-emitted record text.

build_training_compatible_prompt_token_ids

Build the token prefix seen before record continuation during training.

build_partial_record_prefix(*, columns, schema, group_column, group_id, timestamp_column, start_timestamp)

Build an incomplete first record matching the training serialization.

Parameters:

Name Type Description Default
columns Sequence[str]

Saved JSON schema columns in generation order.

required
schema Mapping[str, object]

Saved JSON schema used to coerce known values.

required
group_column str

Configured group column, including the pseudo-group value.

required
group_id object

Group value for this generation stream.

required
timestamp_column str

Configured timestamp column.

required
start_timestamp str | int

First timestamp to generate.

required

Returns:

Type Description
str

An incomplete JSON record ending with the opening quote of the next

str

field name. Including the training ," token keeps the standalone

str

prefix tokenization identical to the beginning of a complete training

str

record. The record begins directly with { because training places

str

the sequence BOS token immediately before the first JSON byte.

Examples:

Given columns beginning with acct_id, txn_index, and cardholder, a group ID of "ACCT-001", and a starting transaction index of 1, the returned prefix is::

{"acct_id":"ACCT-001","txn_index":1,"

The model completes the next field name and the rest of the record.

Raises:

Type Description
GenerationError

If the saved artifact cannot support a partial prefix.

Source code in src/nemo_safe_synthesizer/generation/timeseries_prompting.py
def build_partial_record_prefix(
    *,
    columns: Sequence[str],
    schema: Mapping[str, object],
    group_column: str,
    group_id: object,
    timestamp_column: str,
    start_timestamp: str | int,
) -> str:
    """Build an incomplete first record matching the training serialization.

    Args:
        columns: Saved JSON schema columns in generation order.
        schema: Saved JSON schema used to coerce known values.
        group_column: Configured group column, including the pseudo-group value.
        group_id: Group value for this generation stream.
        timestamp_column: Configured timestamp column.
        start_timestamp: First timestamp to generate.

    Returns:
        An incomplete JSON record ending with the opening quote of the next
        field name. Including the training ``,"`` token keeps the standalone
        prefix tokenization identical to the beginning of a complete training
        record. The record begins directly with ``{`` because training places
        the sequence BOS token immediately before the first JSON byte.

    Examples:
        Given columns beginning with ``acct_id``, ``txn_index``, and
        ``cardholder``, a group ID of ``"ACCT-001"``, and a starting transaction
        index of ``1``, the returned prefix is::

            {"acct_id":"ACCT-001","txn_index":1,"

        The model completes the next field name and the rest of the record.

    Raises:
        GenerationError: If the saved artifact cannot support a partial prefix.
    """
    seed_values: dict[str, object] = {}
    # The pseudo-group is internal bookkeeping for an originally ungrouped
    # dataset, so it must not appear in the generated record prefix.
    if group_column != PSEUDO_GROUP_COLUMN:
        seed_values[group_column] = group_id
    seed_values[timestamp_column] = start_timestamp

    prefix_columns = [column for column in columns if column in seed_values]
    expected_prefix_columns = (
        [timestamp_column] if group_column == PSEUDO_GROUP_COLUMN else [group_column, timestamp_column]
    )
    if (
        prefix_columns != expected_prefix_columns
        or list(columns[: len(expected_prefix_columns)]) != expected_prefix_columns
    ):
        raise GenerationError(
            "The saved time-series schema does not begin with the configured group and timestamp columns. "
            "Retrain the model with the current time-series preprocessing before generating."
        )

    ordered_values = {column: _coerce_prefix_value(schema, column, seed_values[column]) for column in prefix_columns}
    serialized = records_to_jsonl([ordered_values]).rstrip("\n")
    if not serialized.endswith("}"):
        raise GenerationError("Could not serialize the initial time-series partial record.")
    return f'{serialized[:-1]},"'

build_record_history(records)

Build prompt history from exact model-emitted record text.

Training inserts the sequence BOS token immediately before the first record, with no intervening whitespace. The caller adds that BOS token separately, so this helper returns only the exact newline-terminated record bytes.

Parameters:

Name Type Description Default
records Sequence[ParsedRecord]

Accepted records in chronological order.

required

Returns:

Type Description
str

A sequence of newline-terminated records, or an empty string when no

str

records are provided.

Source code in src/nemo_safe_synthesizer/generation/timeseries_prompting.py
def build_record_history(records: Sequence[ParsedRecord]) -> str:
    """Build prompt history from exact model-emitted record text.

    Training inserts the sequence BOS token immediately before the first record,
    with no intervening whitespace. The caller adds that BOS token separately,
    so this helper returns only the exact newline-terminated record bytes.

    Args:
        records: Accepted records in chronological order.

    Returns:
        A sequence of newline-terminated records, or an empty string when no
        records are provided.
    """
    if not records:
        return ""
    return "".join(f"{record.text}\n" for record in records)

build_training_compatible_prompt_token_ids(*, tokenizer, prompt_config, prompt, record_context)

Build the token prefix seen before record continuation during training.

This mirrors Example construction: encode the schema prompt without tokenizer-defined special tokens, apply the configured prompt BOS/EOS tokens, add the sequence BOS token, then append the prefix or history record context. Building IDs explicitly is necessary for tokenizers such as SmolLM3's, which do not add <|im_start|> automatically at inference.

Parameters:

Name Type Description Default
tokenizer EncodeOnlyTokenizer

Tokenizer used by the generation engine.

required
prompt_config LLMPromptConfig

Saved special-token settings.

required
prompt str

Schema prompt text used during training.

required
record_context str | Sequence[str]

Partial first record, or separately encoded history records. Encoding history records individually mirrors training, which tokenizes each newline-terminated JSON record before concatenation.

required

Returns:

Type Description
list[int]

Prompt token IDs whose boundary exactly matches a training example.

Source code in src/nemo_safe_synthesizer/generation/timeseries_prompting.py
def build_training_compatible_prompt_token_ids(
    *,
    tokenizer: EncodeOnlyTokenizer,
    prompt_config: LLMPromptConfig,
    prompt: str,
    record_context: str | Sequence[str],
) -> list[int]:
    """Build the token prefix seen before record continuation during training.

    This mirrors ``Example`` construction: encode the schema prompt without
    tokenizer-defined special tokens, apply the configured prompt BOS/EOS
    tokens, add the sequence BOS token, then append the prefix or history
    record context. Building IDs explicitly is necessary for tokenizers such as
    SmolLM3's, which do not add ``<|im_start|>`` automatically at inference.

    Args:
        tokenizer: Tokenizer used by the generation engine.
        prompt_config: Saved special-token settings.
        prompt: Schema prompt text used during training.
        record_context: Partial first record, or separately encoded history records.
            Encoding history records individually mirrors training, which
            tokenizes each newline-terminated JSON record before concatenation.

    Returns:
        Prompt token IDs whose boundary exactly matches a training example.
    """
    prompt_ids = encode_prompt_token_ids(
        prompt,
        tokenizer=tokenizer,
        prompt_config=prompt_config,
    )

    context_ids: list[int] = []
    context_segments = [record_context] if isinstance(record_context, str) else record_context
    for segment in context_segments:
        context_ids.extend(tokenizer.encode(segment, add_special_tokens=False))
    return [
        *prompt_ids,
        *wrap_sequence_token_ids(
            context_ids,
            prompt_config=prompt_config,
            include_eos=False,
        ),
    ]