Skip to content

Update Deno documentation scraper - #2681

Merged
simon04 merged 5 commits into
freeCodeCamp:mainfrom
srpatcha:chore/add-gitattributes
Sep 13, 2026
Merged

Update Deno documentation scraper#2681
simon04 merged 5 commits into
freeCodeCamp:mainfrom
srpatcha:chore/add-gitattributes

Conversation

@srpatcha

@srpatcha srpatcha commented Apr 25, 2026

Copy link
Copy Markdown
Contributor

No description provided.

Ensure consistent line endings and proper diff handling
for text and binary files.
@srpatcha
srpatcha requested a review from a team as a code owner April 25, 2026 02:42
Add a new UrlScraper for Deno standard library and runtime documentation
(lib/docs/scrapers/deno.rb) with:

- Deno docs site scraping (docs.deno.com)
- Page parsing with Nokogiri (main/article content extraction)
- Link resolution (relative to absolute URL conversion)
- Version handling with semver normalization (v2 and v1 support)
- Module categorization (Web APIs, I/O, File System, Network, etc.)
- Code example extraction with language detection
- HTML filter pipeline (clean_html and entries filters)

Also includes:
- Minitest test class for scraper configuration validation
- Bug fix: replace File.open(path).read with File.read(path) in
  sprites.thor to prevent unclosed file handle leak

Signed-off-by: Srikanth Patchava <spatchava@meta.com>
@simon04 simon04 changed the title chore: add .gitattributes for line ending normalization Update Deno documentation scraper May 26, 2026
@simon04

simon04 commented Jul 19, 2026

Copy link
Copy Markdown
Contributor

The scraper needs fine tuning:

image image

@srpatcha

Copy link
Copy Markdown
Contributor Author

Thanks for the review and the screenshots @simon04! You're right that the output needs refinement.

From the screenshots I can see a few issues to address:

  1. Residual navigation/UI chrome that isn't being stripped by the current clean_html selectors
  2. Code block language detection could be more accurate

I'll iterate on the clean_html.rb and entries.rb filters to clean up the rendering. Could you confirm which Deno docs section/URL the screenshots were captured from so I can target the exact page structure when tuning the selectors? That will help me verify the fixes against the same pages you're seeing.

@simon04 simon04 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

.

Remove source-site navigation and controls that leaked into generated API and runtime pages, while preserving the source language metadata DevDocs uses for syntax highlighting. Add representative regression fixtures for the reviewer-reported page shapes.

Test Plan:
- Ran the focused Deno suite: 17 runs, 73 assertions, 0 failures.
- Ran the remaining Ruby suite excluding two independently failing Windows-only baseline files: 665 runs, 927 assertions, 0 failures.
- Compared old and new filters on current Network and @std/fmt source pages; all measured chrome counts dropped to zero while TypeScript, JavaScript, and shell language metadata was preserved.
- Ran Ruby syntax checks and git diff --check.
@srpatcha

srpatcha commented Sep 8, 2026

Copy link
Copy Markdown
Contributor Author

@simon04 I pushed d17aafb to address the requested scraper fine-tuning without waiting for the missing source URL. I matched the screenshots to the current pages and validated the previous vs updated filter on both:

  • https://docs.deno.com/api/deno/network/: breadcrumbs 1 -> 0, symbol-kind badges 64 -> 0, code-copy buttons 23 -> 0, feedback sections 1 -> 0; all 23 TypeScript blocks still emit pre[data-language="ts"].
  • https://docs.deno.com/runtime/reference/std/fmt/: copy-page controls 1 -> 0, code-copy buttons 2 -> 0, mobile “On this page” blocks 1 -> 0, heading anchors 1 -> 0, feedback sections 1 -> 0; JavaScript and shell blocks remain js and sh.

The filter now scopes output to main#content article, removes the confirmed source-site chrome, and transfers each live language-* class to the enclosing pre[data-language] instead of defaulting every block to TypeScript. I also added reduced API/runtime fixtures and filter/entry regressions.

Validation:

  • Focused Deno tests: 17 runs, 73 assertions, 0 failures/errors.
  • Remaining Ruby suite excluding two independently reproduced Windows-only baseline files: 665 runs, 927 assertions, 0 failures/errors.
  • Full suite reached 761 runs; only the existing Windows case-sensitivity URL assertion and FileStore test/tmp path errors failed when run independently.
  • Ruby syntax checks, editor diagnostics, change validation, and git diff --check passed.
  • The live Thor scraper reached both exact URLs; local Windows curl then stopped at its CA-path configuration, so the before/after output comparison used the already-downloaded current HTML through the real DevDocs parser/filter stack.

Re-review is requested because the two issues identified from the screenshots are now covered directly.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The v1 URL configuration is invalid, API symbols are omitted from the index, and version and styling metadata regress.

Get a fresh assessment by requesting another Copilot review.

Pull request overview

Updates the Deno scraper for the current documentation layout, adds fixtures and tests, and configures line-ending normalization.

Changes:

  • Expands Deno crawling, filtering, and version configuration.
  • Adds scraper and HTML-cleaning tests with fixtures.
  • Adds .gitattributes and simplifies file reading in the sprite task.
File summaries
File Description
.gitattributes Defines text normalization and binary-file handling.
lib/docs/scrapers/deno.rb Reconfigures Deno scraping and versions.
lib/docs/filters/deno/entries.rb Updates entry names and categories.
lib/docs/filters/deno/clean_html.rb Handles the current page structure.
lib/tasks/sprites.thor Uses File.read for ERB input.
test/lib/docs/scrapers/deno_test.rb Adds Deno scraper and filter tests.
test/files/deno_api.html Adds an API-page fixture.
test/files/deno_runtime.html Adds a runtime-page fixture.
Review details
  • Files reviewed: 8/8 changed files
  • Comments generated: 6
  • Review effort level: Balanced

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread lib/docs/scrapers/deno.rb
Comment on lines 45 to 48
version '1' do
self.release = '1.27.0'
self.release = '1.46.3'
self.base_url = 'https://docs.deno.com/api/'
end
Comment on lines 9 to +11
def get_name
if result[:path].start_with?('api/deno/')
at_css('main[id!="content"]')['id'][/\Asymbol_([.\w]+)/, 1]
else
at_css('main article h1').content
end
name = at_css('h1')
name ? name.content.strip : slug.split('/').last
Comment thread lib/docs/scrapers/deno.rb Outdated
class Deno < UrlScraper
self.name = 'Deno'
self.type = 'simple'
self.type = 'deno'
Comment thread lib/docs/scrapers/deno.rb Outdated
/\Aapi\/deno\/~\/Deno\.OpMetrics/, # deprecated in deno 2
]
options[:trailing_slash] = false
self.release = '2.3.1'
Comment thread lib/docs/scrapers/deno.rb

# https://github.com/denoland/manual/blob/main/LICENSE
# https://github.com/denoland/deno/blob/main/LICENSE.md
html_filters.push 'deno/clean_html', 'deno/entries'
Comment thread lib/docs/scrapers/deno.rb Outdated
Comment on lines +87 to +91
def parse_page(response)
doc = Nokogiri::HTML.parse(response.body)
return nil if doc.at_css('meta[http-equiv="refresh"]')

content = doc.at_css('main, article, [role="main"], .markdown-body')

@simon04 simon04 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you!

@simon04
simon04 merged commit ee39489 into freeCodeCamp:main Sep 13, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants