<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Yaml Crawler :: AY的文档</title>
    <link>https://ops.docs.72602.space/zh/demo/design/yaml-crawler/index.html</link>
    <description>Steps define which url you wanna crawl, lets say https://www.xxx.com/aaa.apex create a page pojo to describe what kind of web page you need to process Then you can create a yaml file named root-pages.yaml and its content is&#xA;- &#39;@class&#39;: &#34;org.example.business.hs.code.MainPage&#34; url: &#34;https://www.xxx.com/aaa.apex&#34; and then define a process flow yaml file, implying how to process web pages the crawler will meet. processorChain: - &#39;@class&#39;: &#34;net.zjvis.lab.nebula.crawler.core.processor.decorator.ExceptionRecord&#34; processor: &#39;@class&#39;: &#34;net.zjvis.lab.nebula.crawler.core.processor.decorator.RetryControl&#34; processor: &#39;@class&#39;: &#34;net.zjvis.lab.nebula.crawler.core.processor.decorator.SpeedControl&#34; processor: &#39;@class&#39;: &#34;org.example.business.hs.code.MainPageProcessor&#34; application: &#34;hs-code&#34; time: 100 unit: &#34;MILLISECONDS&#34; retryTimes: 1 - &#39;@class&#39;: &#34;net.zjvis.lab.nebula.crawler.core.processor.decorator.ExceptionRecord&#34; processor: &#39;@class&#39;: &#34;net.zjvis.lab.nebula.crawler.core.processor.decorator.RetryControl&#34; processor: &#39;@class&#39;: &#34;net.zjvis.lab.nebula.crawler.core.processor.decorator.SpeedControl&#34; processor: &#39;@class&#39;: &#34;net.zjvis.lab.nebula.crawler.core.processor.download.DownloadProcessor&#34; pagePersist: &#39;@class&#39;: &#34;org.example.business.hs.code.persist.DownloadPageDatabasePersist&#34; downloadPageRepositoryBeanName: &#34;downloadPageRepository&#34; downloadPageTransformer: &#39;@class&#39;: &#34;net.nebula.crawler.download.DefaultDownloadPageTransformer&#34; skipExists: &#39;@class&#39;: &#34;net.nebula.crawler.download.SkipExistsById&#34; time: 1 unit: &#34;SECONDS&#34; retryTimes: 1 nThreads: 1 pollWaitingTime: 30 pollWaitingTimeUnit: &#34;SECONDS&#34; waitFinishedTimeout: 180 waitFinishedTimeUnit: &#34;SECONDS&#34; ExceptionRecord, RetryControl, SpeedControl are provided by the yaml crawler itself, don’t worry. you only need to extend how to process your page MainPage, for example, you defined a MainPageProcessor. each processor will produce a set of other page or DownloadPage. DownloadPage like a ship containing information you need, and this framework will help you process DownloadPage and download or persist.</description>
    <generator>Hugo</generator>
    <language>zh</language>
    <lastBuildDate></lastBuildDate>
    <atom:link href="https://ops.docs.72602.space/zh/demo/design/yaml-crawler/index.xml" rel="self" type="application/rss+xml" />
  </channel>
</rss>