Skip to content

[Bug]: BFSDeepCrawlStrategy: max_pages is checked per parent page, so a depth level can exceed the limit by a lot #2338

Description

@ipokestuff

crawl4ai version

0.7.4 (seen on a real site), 0.9.4

Expected Behavior

With BFSDeepCrawlStrategy(max_depth=2, max_pages=50) in batch mode, the crawl should stop at 50 pages in total. When the budget is almost used up, link_discovery should only queue as many URLs for the next depth as there is budget left, counting the URLs other parent pages on the same level have already queued.

Current Behavior

The budget is checked per parent page (self.max_pages - self._pages_crawled) and ignores the URLs already in next_level. Every parent on a level can queue up to the full remaining budget, so the next level overshoots the limit.

On a real company website with max_pages=50, the crawl returned 193 pages (1 at depth 0, 35 at depth 1, 157 at depth 2) and took 414 s instead of about 100 s.

Reproduced without a browser on 0.7.4 and 0.9.4: with 14 pages of budget left and 35 parents linking to 30 new pages each, 490 URLs are queued for the next level instead of 14. Script below.

Is this reproducible?

Yes

Inputs Causing the Bug

Steps to Reproduce

Code snippets

OS

Windows

Python version

3.11.9

Browser

No response

Browser version

No response

Error logs & Screenshots (if applicable)

What happens

With max_pages=50 and max_depth=2, a batch deep crawl of a mid-sized company website returned 193 pages: 1 at depth 0, 35 at depth 1 and 157 at depth 2. It took 414 s instead of about 100 s.

Cause

In link_discovery, the remaining budget is calculated as self.max_pages - self._pages_crawled. Links already added to next_level by other parent pages in the same level are not counted. So every parent on depth 1 can queue up to the full remaining budget, and the whole next level gets crawled.

Reproduction (no browser needed)

import asyncio
from types import SimpleNamespace
from crawl4ai.deep_crawling.bfs_strategy import BFSDeepCrawlStrategy

async def main():
    s = BFSDeepCrawlStrategy(max_depth=2, include_external=False, max_pages=50)
    s._pages_crawled = 36  # homepage + 35 depth-1 pages already crawled
    visited, next_level, depths = set(), [], {}
    for i in range(35):
        links = [{"href": f"https://example.com/p{i}/sub{j}"} for j in range(30)]
        r = SimpleNamespace(links={"internal": links, "external": []}, metadata={})
        await s.link_discovery(r, f"https://example.com/p{i}", 1, visited, next_level, depths)
    print("budget left: 14, queued for depth 2:", len(next_level))

asyncio.run(main())

Output on 0.7.4 and 0.9.4: budget left: 14, queued for depth 2: 490

Expected: 14 URLs queued, so the crawl stops at 50 pages.

Suggested fix

Count the queued links in the budget:

remaining_capacity = self.max_pages - self._pages_crawled - len(next_level)

As a workaround I subclassed the strategy and added len(next_level) to _pages_crawled for the duration of link_discovery. With that, the same site returns exactly 50 pages (1 / 35 / 14 by depth) in 99 s.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    🐞 BugSomething isn't working🩺 Needs TriageNeeds attention of maintainers

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions