crawl4ai version
0.7.4 (seen on a real site), 0.9.4
Expected Behavior
With BFSDeepCrawlStrategy(max_depth=2, max_pages=50) in batch mode, the crawl should stop at 50 pages in total. When the budget is almost used up, link_discovery should only queue as many URLs for the next depth as there is budget left, counting the URLs other parent pages on the same level have already queued.
Current Behavior
The budget is checked per parent page (self.max_pages - self._pages_crawled) and ignores the URLs already in next_level. Every parent on a level can queue up to the full remaining budget, so the next level overshoots the limit.
On a real company website with max_pages=50, the crawl returned 193 pages (1 at depth 0, 35 at depth 1, 157 at depth 2) and took 414 s instead of about 100 s.
Reproduced without a browser on 0.7.4 and 0.9.4: with 14 pages of budget left and 35 parents linking to 30 new pages each, 490 URLs are queued for the next level instead of 14. Script below.
Is this reproducible?
Yes
Inputs Causing the Bug
Steps to Reproduce
Code snippets
OS
Windows
Python version
3.11.9
Browser
No response
Browser version
No response
Error logs & Screenshots (if applicable)
What happens
With max_pages=50 and max_depth=2, a batch deep crawl of a mid-sized company website returned 193 pages: 1 at depth 0, 35 at depth 1 and 157 at depth 2. It took 414 s instead of about 100 s.
Cause
In link_discovery, the remaining budget is calculated as self.max_pages - self._pages_crawled. Links already added to next_level by other parent pages in the same level are not counted. So every parent on depth 1 can queue up to the full remaining budget, and the whole next level gets crawled.
Reproduction (no browser needed)
import asyncio
from types import SimpleNamespace
from crawl4ai.deep_crawling.bfs_strategy import BFSDeepCrawlStrategy
async def main():
s = BFSDeepCrawlStrategy(max_depth=2, include_external=False, max_pages=50)
s._pages_crawled = 36 # homepage + 35 depth-1 pages already crawled
visited, next_level, depths = set(), [], {}
for i in range(35):
links = [{"href": f"https://example.com/p{i}/sub{j}"} for j in range(30)]
r = SimpleNamespace(links={"internal": links, "external": []}, metadata={})
await s.link_discovery(r, f"https://example.com/p{i}", 1, visited, next_level, depths)
print("budget left: 14, queued for depth 2:", len(next_level))
asyncio.run(main())
Output on 0.7.4 and 0.9.4: budget left: 14, queued for depth 2: 490
Expected: 14 URLs queued, so the crawl stops at 50 pages.
Suggested fix
Count the queued links in the budget:
remaining_capacity = self.max_pages - self._pages_crawled - len(next_level)
As a workaround I subclassed the strategy and added len(next_level) to _pages_crawled for the duration of link_discovery. With that, the same site returns exactly 50 pages (1 / 35 / 14 by depth) in 99 s.
crawl4ai version
0.7.4 (seen on a real site), 0.9.4
Expected Behavior
With
BFSDeepCrawlStrategy(max_depth=2, max_pages=50)in batch mode, the crawl should stop at 50 pages in total. When the budget is almost used up,link_discoveryshould only queue as many URLs for the next depth as there is budget left, counting the URLs other parent pages on the same level have already queued.Current Behavior
The budget is checked per parent page (
self.max_pages - self._pages_crawled) and ignores the URLs already innext_level. Every parent on a level can queue up to the full remaining budget, so the next level overshoots the limit.On a real company website with
max_pages=50, the crawl returned 193 pages (1 at depth 0, 35 at depth 1, 157 at depth 2) and took 414 s instead of about 100 s.Reproduced without a browser on 0.7.4 and 0.9.4: with 14 pages of budget left and 35 parents linking to 30 new pages each, 490 URLs are queued for the next level instead of 14. Script below.
Is this reproducible?
Yes
Inputs Causing the Bug
Steps to Reproduce
Code snippets
OS
Windows
Python version
3.11.9
Browser
No response
Browser version
No response
Error logs & Screenshots (if applicable)
What happens
With
max_pages=50andmax_depth=2, a batch deep crawl of a mid-sized company website returned 193 pages: 1 at depth 0, 35 at depth 1 and 157 at depth 2. It took 414 s instead of about 100 s.Cause
In
link_discovery, the remaining budget is calculated asself.max_pages - self._pages_crawled. Links already added tonext_levelby other parent pages in the same level are not counted. So every parent on depth 1 can queue up to the full remaining budget, and the whole next level gets crawled.Reproduction (no browser needed)
Output on 0.7.4 and 0.9.4:
budget left: 14, queued for depth 2: 490Expected: 14 URLs queued, so the crawl stops at 50 pages.
Suggested fix
Count the queued links in the budget:
As a workaround I subclassed the strategy and added
len(next_level)to_pages_crawledfor the duration oflink_discovery. With that, the same site returns exactly 50 pages (1 / 35 / 14 by depth) in 99 s.