<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Auto Scaling :: Ay Docs</title>
    <link>https://ops.docs.72602.space/kubernetes/serverless/kserve/auto_scaling/index.html</link>
    <description>Soft Limit You can configure InferenceService with annotation autoscaling.knative.dev/target for a soft limit. The soft limit is a targeted limit rather than a strictly enforced bound, particularly if there is a sudden burst of requests, this value can be exceeded.&#xA;apiVersion: &#34;serving.kserve.io/v1beta1&#34; kind: &#34;InferenceService&#34; metadata: name: &#34;sklearn-iris&#34; namespace: kserve-test annotations: autoscaling.knative.dev/target: &#34;5&#34; spec: predictor: model: args: [&#34;--enable_docs_url=True&#34;] modelFormat: name: sklearn resources: {} runtime: kserve-sklearnserver storageUri: &#34;gs://kfserving-examples/models/sklearn/1.0/model&#34; Hard Limit You can also configure InferenceService with field containerConcurrency with a hard limit. The hard limit is an enforced upper bound. If concurrency reaches the hard limit, surplus requests will be buffered and must wait until enough capacity is free to execute the requests.</description>
    <generator>Hugo</generator>
    <language>en</language>
    <lastBuildDate></lastBuildDate>
    <atom:link href="https://ops.docs.72602.space/kubernetes/serverless/kserve/auto_scaling/index.xml" rel="self" type="application/rss+xml" />
  </channel>
</rss>