Skip to content

feat: add cognee-community-tasks-firecrawl - #214

Open
JuampiHernandez wants to merge 1 commit into
topoteretes:mainfrom
JuampiHernandez:feat/firecrawl-tasks
Open

JuampiHernandez wants to merge 1 commit into
topoteretes:mainfrom
JuampiHernandez:feat/firecrawl-tasks

Conversation

@JuampiHernandez

Copy link
Copy Markdown

Hi, I'm Juan from the Firecrawl team.

Summary

This adds cognee-community-tasks-firecrawl under packages/task/firecrawl_tasks, with four async tasks, scrape_urls, scrape_and_add, search_web and search_and_add, plus one row for it in the root README package table.

Scrape turns a URL, including JavaScript-rendered pages and PDFs, into markdown for cognee to chunk and graph. Search returns each result page's markdown in the same call, so search_and_add does not need a second request per result.

Why

This repo already uses Firecrawl. The dlt and Qdrant docs assistants in experimental/ scrape their docs with hand-written requests calls to the old v1 scrape endpoint, with the key set as a constant at the top of each script. This package turns that step into a tested cognee task, so the same flow becomes scrape_and_add(urls) with the key read from the environment, on the current v2 API. Search adds the case where you start from a question instead of a list of URLs.

Package

It follows the shape of the existing exa_tasks package, with a typed result dataclass, a client builder that reads the key from the environment, cognee.add then cognee.cognify, mocked tests, an example and a README. It uses the async client from the official firecrawl-py SDK and requires FIRECRAWL_API_KEY (or an api_key argument). cognee is pinned to 1.6.1, the version #210 moved the adapters to.

Example

import asyncio
import cognee
from cognee_community_tasks_firecrawl import scrape_and_add, search_and_add

async def main():
    await scrape_and_add(["https://docs.cognee.ai/"], dataset_name="firecrawl")
    await search_and_add("How do knowledge graphs improve LLM memory?", limit=3)
    print(await cognee.search("What is cognee?"))

asyncio.run(main())

scrape_urls runs pages concurrently under a concurrency cap (default 5) and keeps input order. A URL that fails does not stop the batch, it comes back with empty content and the SDK's error in error. An invalid key or exhausted credits raises the SDK error instead, because it would fail every URL. The *_and_add tasks skip pages with no markdown or an HTTP error status, and cognify only the dataset they wrote to.

Tested

All of this ran inside packages/task/firecrawl_tasks.

  • uv sync --all-extras
  • ruff check . and ruff format --check . are clean with ruff 0.16.2 and 0.16.9 (the version the last ruff workflow run installed)
  • uv run --with pytest pytest tests -q gives 19 passed on Python 3.11, 3.12 and 3.13, with AsyncFirecrawl and cognee mocked and fixtures built from the SDK's own Document and SearchData types
  • With real keys in a Linux container, uv run python ./examples/example.py ran to completion, and a longer live run scraped the cognee docs and the Attention Is All You Need PDF (3.6k and 44.5k chars of markdown) and ran search_and_add with 3 results. cognee.search answered from both datasets. A bad key raised the SDK's UnauthorizedError and a missing key raised a ValueError naming FIRECRAWL_API_KEY

I did not add a job to community_task_tests.yml, since it would need a Firecrawl secret in the repo. Happy to add one if you want it.

If you would rather start with a scrape-only package, I can drop the two search tasks from this PR.

I affirm that all code in every commit of this pull request conforms to the terms of the Topoteretes Developer Certificate of Origin.

Signed-off-by: Juampi <JuampiHernandez@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant