Search in book...
Toggle Font Controls
Create new playlist

Name your new playlist

Playlist description (optional)
Sign In

Email address

Password

Forgot Password?

or

Continue with Facebook

Continue with Google
Sign Up

Full Name

Email address

Confirm Email Address

Password

or

Continue with Facebook

Continue with Google

Concurrent Downloading

In the previous chapters, our crawlers downloaded web pages sequentially, waiting for each download to complete before starting the next one. Sequential downloading is fine for the relatively small example website but quickly becomes impractical for larger crawls. To crawl a large website of one million web pages at an average of one web page per second would take over 11 days of continuous downloading. This time can be significantly improved by downloading multiple web pages simultaneously.

This chapter will cover downloading web pages with multiple threads and processes and comparing the performance with sequential downloading.

In this chapter, we will cover the following topics:

One million web pages
Sequential crawler
Threaded crawler
Multiprocessing crawler

..................Content has been hidden....................

You can't read the all page of ebook, please click here login for view all page.

Table of Contents for Concurrent Downloading

Create new playlist

Sign In

Sign Up

Table of Contents for
Concurrent Downloading