PHP cURL crawler doesn't fetch all data

Question

I'm trying to write my first crawler by using PHP with cURL library. My aim is to fetch data from one site systematically, which means that the code doesn't follow all hyperlinks on the given site but only specific links.

Logic of my code is to go to the main page and get links for several categories and store those in an array. Once it's done the crawler goes to those category sites on the page and looks if the category has more than one pages. If so, it stores subpages also in another array. Finally I merge the arrays to get all the links for sites that needs to be crawled and start to fetch required data.

I call the below function to start a cURL session and fetch data to a variable, which I pass to a DOM object later and parse it with Xpath. I store cURL total_time and http_code in a log file.

The problem is that the crawler runs for 5-6 minutes then stops and doesn't fetch all required links for sub-pages. I print content of arrays to check result. I can't see any http error in my log, all sites give a http 200 status code. I can't see any PHP related error even if I turn on PHP debug on my localhost.

I assume that the site blocks my crawler after few minutes because of too many requests but I'm not sure. Is there any way to get a more detailed debug? Do you think that PHP is adequate for this type of activity because I wan't to use the same mechanism to fetch content from more than 100 other sites later on?

My cURL code is as follows:

function get_url($url)
{
    $ch = curl_init();
    curl_setopt($ch, CURLOPT_HEADER, 0);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, 1);
    curl_setopt($ch, CURLOPT_CONNECTTIMEOUT, 30);
    curl_setopt($ch, CURLOPT_URL, $url);
    $data = curl_exec($ch);
    $info = curl_getinfo($ch);  
    $logfile = fopen("crawler.log","a");
    echo fwrite($logfile,'Page ' . $info['url'] . ' fetched in ' . $info['total_time'] . ' seconds. Http status code: ' . $info['http_code'] . "\n");
    fclose($logfile);
    curl_close($ch);

    return $data;
}

// Start to crawle main page.

$site2crawl = 'http://www.site.com/';

$dom = new DOMDocument();
@$dom->loadHTML(get_url($site2crawl));
$xpath = new DomXpath($dom);

I found this line in my LAMPP erro_log: [:error] [pid 2996] [client 127.0.0.1:49848] PHP Fatal error: Maximum execution time of 30 seconds exceeded in /opt/lampp/htdocs/clw/clw.php on line 73. I'll try to increase timeout for cURL and retry. — g0m3z
– g0m3z, Commented Dec 31, 2012 at 19:58
I increased the timeout parameter then changed to zero but it did not help. — g0m3z
– g0m3z, Commented Dec 31, 2012 at 20:11
Have you seen if curl is getting any errors? Something like this should work: if( $data == false ) { fwrite( $logfile, curl_error( $ch ); ) } — Chris Ostmo
– Chris Ostmo, Commented Dec 31, 2012 at 20:42
By 'increase timeout for cURL' do you mean you used set_time_limit? — Quentin Skousen
– Quentin Skousen, Commented Dec 31, 2012 at 20:47
Thanks to kkhugs who suggested to set the time limit to zero within the code. It helped. The following code solved my issue: set_time_limit(0); I also implemented the code which can be found here to avoid memory leak issue. Thread can be closed. Thanks for everyone! gomez — g0m3z
– g0m3z, Commented Dec 31, 2012 at 23:25

Quentin Skousen · Accepted Answer · 2013-01-01 12:12:15Z

1

Use set_time_limit to extend the amount of time your script can run for. That is why you are getting Fatal error: Maximum execution time of 30 seconds exceeded in your error log.

answered Jan 1, 2013 at 12:12

Quentin Skousen

1,0572 gold badges18 silver badges31 bronze badges

Sign up to request clarification or add additional context in comments.

Comments

user1938139 · Accepted Answer · 2012-12-31 21:15:52Z

0

do you need to run this on a server? If not, you should try the cli version of php - it is exempt from common restrictions

answered Dec 31, 2012 at 21:15

user1938139

1411 silver badge5 bronze badges

3 Comments

g0m3z Over a year ago

Yes, I would like to run it on a server later on in production.

Toby Allen Over a year ago

why would you not be able to run the cli version on a server?

g0m3z Over a year ago

Thanks @TobyAllen my issue was solved already. I'll have enough time to figure out later how will I implement this in production. I'm going to improve my crawler code first (with parallel threads for example).

Collectives™ on Stack Overflow

PHP cURL crawler doesn't fetch all data

2 Answers 2

Comments

3 Comments

Your Answer

Linked

Hot Network Questions

Collectives™ on Stack Overflow

2 Answers 2

Comments

3 Comments

Your Answer

Sign up or log in

Post as a guest

Linked

Related