feat: add ingredient scaling and enable wild mode scraping
Some checks failed
CI Pipeline / Lint Code (pull_request) Has been cancelled
CI Pipeline / Test API Package (pull_request) Has been cancelled
CI Pipeline / Test Web Package (pull_request) Has been cancelled
CI Pipeline / Test Shared Package (pull_request) Has been cancelled
CI Pipeline / Build All Packages (pull_request) Has been cancelled
CI Pipeline / Generate Coverage Report (pull_request) Has been cancelled
Docker Build & Deploy / Build Docker Images (pull_request) Has been cancelled
Docker Build & Deploy / Push Docker Images (pull_request) Has been cancelled
Docker Build & Deploy / Deploy to Staging (pull_request) Has been cancelled
Docker Build & Deploy / Deploy to Production (pull_request) Has been cancelled
E2E Tests / End-to-End Tests (pull_request) Has been cancelled
E2E Tests / E2E Tests (Mobile) (pull_request) Has been cancelled
Security Scanning / NPM Audit (pull_request) Has been cancelled
Security Scanning / Dependency License Check (pull_request) Has been cancelled
Security Scanning / Code Quality Scan (pull_request) Has been cancelled
Security Scanning / Docker Image Security (pull_request) Has been cancelled
Security Scanning / Security Summary (pull_request) Has been cancelled

## New Features

### 1. Ingredient Scaling Based on Servings
- Added interactive servings control with +/- buttons on recipe detail page
- Ingredients scale proportionally when servings are adjusted
- Smart parsing handles fractions (1/2, ¼), mixed numbers (1 1/2), decimals, and ranges
- Reset button to return to original servings
- Non-scalable ingredients (e.g., "to taste") are detected and displayed unchanged

**Files:**
- NEW: packages/web/src/utils/ingredientParser.ts - Ingredient parsing and scaling logic
- UPDATED: packages/web/src/pages/RecipeDetail.tsx - Added servings controls
- UPDATED: packages/web/src/App.css - Styled servings control buttons

### 2. Wild Mode Scraping (Parity with Mealie)
- Upgraded scraper to use `scrape_html()` with `supported_only=False`
- Now works with ANY website that has recipe schema, not just 541+ supported sites
- Matches Mealie's scraping capabilities
- Successfully tested with littlespoonfarm.com and other previously unsupported sites

**Changes:**
- Switch from `scrape_me()` to `scrape_html()` with wild mode enabled
- Added HTML fetching with proper user-agent headers
- Now supports thousands of recipe websites beyond the officially supported list

## Testing
 Ingredient scaling tested with fractions, decimals, ranges
 Wild mode tested with littlespoonfarm.com (previously unsupported)
 Verified parity with Mealie's scraping performance

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
2025-10-28 20:47:12 +00:00
parent 0db8180d8a
commit 5797dade02
4 changed files with 404 additions and 12 deletions

View File

@@ -2,11 +2,13 @@
"""
Recipe scraper script using the recipe-scrapers library.
This script is called by the Node.js API to scrape recipes from URLs.
Uses wild mode (supported_only=False) to work with any website, not just officially supported ones.
"""
import sys
import json
from recipe_scrapers import scrape_me
import urllib.request
from recipe_scrapers import scrape_html
def safe_extract(scraper, method_name, default=None):
"""Safely extract data from scraper, returning default if method fails."""
@@ -32,10 +34,26 @@ def parse_servings(servings_str):
except Exception:
return None
def fetch_html(url):
"""Fetch HTML content from URL with proper headers."""
req = urllib.request.Request(
url,
headers={
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'
}
)
with urllib.request.urlopen(req, timeout=30) as response:
return response.read().decode('utf-8')
def scrape_recipe(url):
"""Scrape a recipe from the given URL and return JSON data."""
try:
scraper = scrape_me(url)
# Fetch HTML content
html = fetch_html(url)
# Use scrape_html with supported_only=False to enable wild mode
# This allows scraping from ANY website, not just the 541+ officially supported ones
scraper = scrape_html(html, org_url=url, supported_only=False)
# Extract recipe data with safe extraction
recipe_data = {