pub struct CallModel {
pub algorithm: String,
pub request: Request,
pub models: Vec<ModelId>,
pub recover_errors: bool,
/* private fields */
}Expand description
An offloaded model call, surfaced inside Step::CallModel.
The host reads the public fields and passes its model-call future to
respond, unblocking the algorithm’s Driver::call_model on the
other side. switchyard-llm-client’s run is the ready-made consumer that does this for
you.
Driver::call_model stamps the first candidate model onto the request before publishing
the call. A consumer that falls through to a later candidate must re-stamp it.
Fields§
§algorithm: StringThe name of the algorithm that produced this call, so a host instrumenting the calls it serves can attribute its own spans to the algorithm behind them.
request: RequestThe request to serve; its model is stamped with the first candidate.
models: Vec<ModelId>Candidate models, tried in order until one answers. Never empty.
recover_errors: boolReturn client errors to the algorithm so it can apply its fallback policy.
Implementations§
Source§impl CallModel
impl CallModel
Sourcepub async fn respond(
self,
work: impl Future<Output = Result<Response>>,
) -> Result<()>
pub async fn respond( self, work: impl Future<Output = Result<Response>>, ) -> Result<()>
Run host work until it completes or the algorithm stops waiting for this call.
Dropping the waiting Driver::call_model future drops work and returns Ok(()).
Errors stop the host’s run unless Self::recover_errors is enabled.
Use std::future::ready(result) for an already available response.
A delivered response stream owns its subsequent lifetime.