Skip to content

Optional source-rights evidence hook/metadata for crawl and scrape results #2345

Description

@reflectme-source

Crawl4AI is directly in the path where this evidence is most useful: public URL -> fetched/extracted content -> RAG/agent/data pipeline. The library already exposes custom strategies/hooks, and the Cloud API/MCP path returns structured page results.

Would you consider an optional provider-neutral enrichment hook that attaches a signed source-rights observation to each page result?

The external provider would receive URL + intended purpose and return timestamped evidence of machine-readable source-rights/TDM declarations. The crawler would not make a legal decision; the receipt would simply travel with result metadata so downstream RAG/agent policy can decide what to do.

I’m building AcqPath as one provider, but the useful abstraction is generic: rights_evidence_provider -> normalized receipt/digest in crawl result metadata.

If that direction fits, I can sketch a minimal hook/result-field contract that works for both self-hosted Crawl4AI and Cloud without making AcqPath a required dependency.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions