5.6 KiB
5.6 KiB
🔧 Bygma Scraper Forbedringer
Dato: 29. Oktober 2025
Status: ✅ Opdateret og klar til test
📝 Problem
Den originale scraper fandt INGEN montagevejledninger selv for produkter vi ved har dem:
❌ Eternit tagplader - INGEN MANUAL
❌ Tagpap - INGEN MANUAL
❌ Tagsten - INGEN MANUAL
❌ Velux vinduer - INGEN MANUAL
❌ Lindab tagrender - INGEN MANUAL (selv dem i database!)
🔍 Root Cause Analysis
Original Kode (Fejl):
for link in links:
title = await link.get_attribute('title')
if title and 'montage' in title.lower(): # ❌ KUN title attribute
# ... rest
Problem:
- Checker KUN
titleattribute - Mange links har "montage" i text content men IKKE i title
- Miss: "Montagevejledning_B7.pdf" (text) vs title=""
Opdateret Kode (Fix):
for link in links:
title = await link.get_attribute('title')
text = await link.inner_text() # ✅ Hent text content
is_montage = False
if title and 'montage' in title.lower():
is_montage = True
if text and 'montage' in text.lower() and '.pdf' in text.lower(): # ✅ Check text
is_montage = True
if is_montage:
# ... rest
Fordele:
- ✅ Checker BÅDE title OG text content
- ✅ Kræver både "montage" og ".pdf" i text (reducer false positives)
- ✅ Forbedret regex til at extract filename
🎯 Ændringer
1. check_manual_availability() - Quick Check
Linje ~165-205
Før:
if title and 'montage' in title.lower():
# ... extract URL
Efter:
title = await link.get_attribute('title')
text = await link.inner_text()
is_montage = False
if title and 'montage' in title.lower():
is_montage = True
if text and 'montage' in text.lower() and '.pdf' in text.lower():
is_montage = True
if is_montage:
# ... extract URL
filename_match = re.search(r'([^/]+\.pdf)', text, re.IGNORECASE)
filename = filename_match.group(1) if filename_match else os.path.basename(pdf_url)
2. scrape_and_analyze() - Full Scrape
Linje ~265-305
Før:
if title and 'montage' in title.lower():
filename = os.path.basename(pdf_url)
Efter:
title = await link.get_attribute('title')
text = await link.inner_text()
is_montage = False
if title and 'montage' in title.lower():
is_montage = True
if text and 'montage' in text.lower() and '.pdf' in text.lower():
is_montage = True
if is_montage:
filename_match = re.search(r'([^/]+\.pdf)', text, re.IGNORECASE)
filename = filename_match.group(1) if filename_match else os.path.basename(pdf_url)
📊 Forventet Impact
Før Opdatering:
- Metode: Kun title attribute check
- Success Rate: ~0-5% (næsten ingen manuals fundet)
- False Negatives: Høj (misser de fleste manuals)
Efter Opdatering:
- Metode: Title + text content check
- Forventet Success Rate: ~60-80%
- False Negatives: Lav (finder de fleste manuals)
- False Positives: Minimal (kræver både "montage" og ".pdf")
🧪 Test Strategi
Fase 1: Test med kendte produkter
# Produkter vi VED har manuals i databasen:
1. Lindab Rainline (077425) - Tagrende
2. Rockwool Vintermåtte (001683) - Isolering
3. Knauf Ecoblanket (106243) - Isolering
4. SWEDOOR SNAP-IN (118266) - Dør
Fase 2: Test tag materialer
# SwissPearl B7 tagplader (brugt i original scraper)
https://www.bygma.dk/proff/byggemateriale/tag/bolgeplader/bolgeplader/swisspearl-b7-tagplader-i-sortbla---hjornehul---1100x570mm200p147072/
Fase 3: Bulk test
# Test 50 tilfældige Bygma produkter og se coverage
python3 bulk_test_manuals.py
🚀 Next Steps
1. Performance Optimization
Problem: Scraper timeout efter 60-120 sekunder
Løsninger:
- Reducer wait timeouts (5000ms → 2000ms)
- Skip cookie accept hvis allerede accepteret
- Cache login session (re-use samme browser context)
- Parallel processing af multiple produkter
2. Robustness
- Retry logic (3 forsøg ved fejl)
- Fallback til alternative selectors
- Better error messages (log hvilken step der fejler)
3. Coverage
- Test 100 mest brugte Bygma produkter
- Identificér produktkategorier med høj manual coverage
- Build prioriteret scraping queue
📁 Filer Opdateret
-
scrape_bygma_api.py - Python scraper
check_manual_availability()- Linje ~165-205scrape_and_analyze()- Linje ~265-305
-
unified-server.js - API endpoint
POST /api/installation-manuals/check-availability- Linje ~4058
-
MANUAL_CHECK_ENHANCEMENT.md - Feature documentation
🎯 Success Metrics
Mål:
- Find manuals for mindst 50% af tag materialer
- Reduce scraping tid til under 30 sekunder per produkt
- Zero false positives (ingen datablade eller sikkerhedsblade)
Status:
- ✅ Kode opdateret
- ⏳ Test pending (timeout issues)
- ⏳ Deployment pending
💡 Observations
Bygma HTML Struktur:
<a ng-click="downloadItem('some-guid', '/files/path/to/Montagevejledning_B7.pdf')">
Montagevejledning_B7.pdf <!-- TEXT CONTENT (vigtig!) -->
</a>
Key Insight: PDF filnavnet er ofte i link text, ikke i title attribute!
Alternative Approaches:
- Playwright Network Interception - Capture download URLs
- JavaScript Evaluation - Execute ng-click og capture PDF URL
- Bygma API - Hvis de har en dokumentation API (unlikely)
Konklusion: Scraper er nu mere robust men har performance issues. Næste step er at optimere hastighed og teste med rigtige produkter.