<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Kserve :: Ay Docs</title>
    <link>https://ops.docs.72602.space/kubernetes/serverless/kserve/index.html</link>
    <description>Install Kserve Serving Inference First Pytorch ISVC First Custom Model First Model In Minio Kafka Sink Transformer Generative First Generative Service Canary Policy Rollout Example Auto Scaling</description>
    <generator>Hugo</generator>
    <language>en</language>
    <lastBuildDate>Thu, 07 Mar 2024 15:00:59 +0800</lastBuildDate>
    <atom:link href="https://ops.docs.72602.space/kubernetes/serverless/kserve/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Install Kserve</title>
      <link>https://ops.docs.72602.space/kubernetes/serverless/kserve/install/index.html</link>
      <pubDate>Thu, 07 Mar 2024 15:00:59 +0800</pubDate>
      <guid>https://ops.docs.72602.space/kubernetes/serverless/kserve/install/index.html</guid>
      <description>Preliminary v 1.30 + Kubernetes has installed, if not check 🔗link Helm has installed, if not check 🔗link Installation Install By Shell Steps ArgoCD Preliminary 1. Kubernetes has installed, if not check 🔗link 2. Helm binary has installed, if not check 🔗link 1.install from script directly</description>
    </item>
    <item>
      <title>Serving</title>
      <link>https://ops.docs.72602.space/kubernetes/serverless/kserve/serving/index.html</link>
      <pubDate>Thu, 07 Mar 2024 15:00:59 +0800</pubDate>
      <guid>https://ops.docs.72602.space/kubernetes/serverless/kserve/serving/index.html</guid>
      <description>Inference First Pytorch ISVC First Custom Model First Model In Minio Kafka Sink Transformer Generative First Generative Service</description>
    </item>
    <item>
      <title>Canary Policy</title>
      <link>https://ops.docs.72602.space/kubernetes/serverless/kserve/canary/index.html</link>
      <pubDate>Thu, 07 Mar 2024 15:00:59 +0800</pubDate>
      <guid>https://ops.docs.72602.space/kubernetes/serverless/kserve/canary/index.html</guid>
      <description>KServe supports canary rollouts for inference services. Canary rollouts allow for a new version of an InferenceService to receive a percentage of traffic. Kserve supports a configurable canary rollout strategy with multiple steps. The rollout strategy can also be implemented to rollback to the previous revision if a rollout step fails.&#xA;KServe automatically tracks the last good revision that was rolled out with 100% traffic. The canaryTrafficPercent field in the component’s spec needs to be set with the percentage of traffic that should be routed to the new revision. KServe will then automatically split the traffic between the last good revision and the revision that is currently being rolled out according to the canaryTrafficPercent value.</description>
    </item>
    <item>
      <title>Auto Scaling</title>
      <link>https://ops.docs.72602.space/kubernetes/serverless/kserve/auto_scaling/index.html</link>
      <pubDate>Thu, 07 Mar 2024 15:00:59 +0800</pubDate>
      <guid>https://ops.docs.72602.space/kubernetes/serverless/kserve/auto_scaling/index.html</guid>
      <description>Soft Limit You can configure InferenceService with annotation autoscaling.knative.dev/target for a soft limit. The soft limit is a targeted limit rather than a strictly enforced bound, particularly if there is a sudden burst of requests, this value can be exceeded.&#xA;apiVersion: &#34;serving.kserve.io/v1beta1&#34; kind: &#34;InferenceService&#34; metadata: name: &#34;sklearn-iris&#34; namespace: kserve-test annotations: autoscaling.knative.dev/target: &#34;5&#34; spec: predictor: model: args: [&#34;--enable_docs_url=True&#34;] modelFormat: name: sklearn resources: {} runtime: kserve-sklearnserver storageUri: &#34;gs://kfserving-examples/models/sklearn/1.0/model&#34; Hard Limit You can also configure InferenceService with field containerConcurrency with a hard limit. The hard limit is an enforced upper bound. If concurrency reaches the hard limit, surplus requests will be buffered and must wait until enough capacity is free to execute the requests.</description>
    </item>
  </channel>
</rss>