Files
tilbudgivern/docs/archived/SCRAPER_IMPROVEMENTS.md
T
2025-10-30 18:28:58 +00:00

5.6 KiB

🔧 Bygma Scraper Forbedringer

Dato: 29. Oktober 2025
Status: ✅ Opdateret og klar til test


📝 Problem

Den originale scraper fandt INGEN montagevejledninger selv for produkter vi ved har dem:

❌ Eternit tagplader - INGEN MANUAL
❌ Tagpap - INGEN MANUAL  
❌ Tagsten - INGEN MANUAL
❌ Velux vinduer - INGEN MANUAL
❌ Lindab tagrender - INGEN MANUAL (selv dem i database!)

🔍 Root Cause Analysis

Original Kode (Fejl):

for link in links:
    title = await link.get_attribute('title')
    if title and 'montage' in title.lower():  # ❌ KUN title attribute
        # ... rest

Problem:

  • Checker KUN title attribute
  • Mange links har "montage" i text content men IKKE i title
  • Miss: "Montagevejledning_B7.pdf" (text) vs title=""

Opdateret Kode (Fix):

for link in links:
    title = await link.get_attribute('title')
    text = await link.inner_text()  # ✅ Hent text content
    
    is_montage = False
    if title and 'montage' in title.lower():
        is_montage = True
    if text and 'montage' in text.lower() and '.pdf' in text.lower():  # ✅ Check text
        is_montage = True
    
    if is_montage:
        # ... rest

Fordele:

  • ✅ Checker BÅDE title OG text content
  • ✅ Kræver både "montage" og ".pdf" i text (reducer false positives)
  • ✅ Forbedret regex til at extract filename

🎯 Ændringer

1. check_manual_availability() - Quick Check

Linje ~165-205

Før:

if title and 'montage' in title.lower():
    # ... extract URL

Efter:

title = await link.get_attribute('title')
text = await link.inner_text()

is_montage = False
if title and 'montage' in title.lower():
    is_montage = True
if text and 'montage' in text.lower() and '.pdf' in text.lower():
    is_montage = True

if is_montage:
    # ... extract URL
    filename_match = re.search(r'([^/]+\.pdf)', text, re.IGNORECASE)
    filename = filename_match.group(1) if filename_match else os.path.basename(pdf_url)

2. scrape_and_analyze() - Full Scrape

Linje ~265-305

Før:

if title and 'montage' in title.lower():
    filename = os.path.basename(pdf_url)

Efter:

title = await link.get_attribute('title')
text = await link.inner_text()

is_montage = False
if title and 'montage' in title.lower():
    is_montage = True
if text and 'montage' in text.lower() and '.pdf' in text.lower():
    is_montage = True

if is_montage:
    filename_match = re.search(r'([^/]+\.pdf)', text, re.IGNORECASE)
    filename = filename_match.group(1) if filename_match else os.path.basename(pdf_url)

📊 Forventet Impact

Før Opdatering:

  • Metode: Kun title attribute check
  • Success Rate: ~0-5% (næsten ingen manuals fundet)
  • False Negatives: Høj (misser de fleste manuals)

Efter Opdatering:

  • Metode: Title + text content check
  • Forventet Success Rate: ~60-80%
  • False Negatives: Lav (finder de fleste manuals)
  • False Positives: Minimal (kræver både "montage" og ".pdf")

🧪 Test Strategi

Fase 1: Test med kendte produkter

# Produkter vi VED har manuals i databasen:
1. Lindab Rainline (077425) - Tagrende
2. Rockwool Vintermåtte (001683) - Isolering
3. Knauf Ecoblanket (106243) - Isolering
4. SWEDOOR SNAP-IN (118266) - Dør

Fase 2: Test tag materialer

# SwissPearl B7 tagplader (brugt i original scraper)
https://www.bygma.dk/proff/byggemateriale/tag/bolgeplader/bolgeplader/swisspearl-b7-tagplader-i-sortbla---hjornehul---1100x570mm200p147072/

Fase 3: Bulk test

# Test 50 tilfældige Bygma produkter og se coverage
python3 bulk_test_manuals.py

🚀 Next Steps

1. Performance Optimization

Problem: Scraper timeout efter 60-120 sekunder

Løsninger:

  • Reducer wait timeouts (5000ms → 2000ms)
  • Skip cookie accept hvis allerede accepteret
  • Cache login session (re-use samme browser context)
  • Parallel processing af multiple produkter

2. Robustness

  • Retry logic (3 forsøg ved fejl)
  • Fallback til alternative selectors
  • Better error messages (log hvilken step der fejler)

3. Coverage

  • Test 100 mest brugte Bygma produkter
  • Identificér produktkategorier med høj manual coverage
  • Build prioriteret scraping queue

📁 Filer Opdateret

  1. scrape_bygma_api.py - Python scraper

    • check_manual_availability() - Linje ~165-205
    • scrape_and_analyze() - Linje ~265-305
  2. unified-server.js - API endpoint

    • POST /api/installation-manuals/check-availability - Linje ~4058
  3. MANUAL_CHECK_ENHANCEMENT.md - Feature documentation


🎯 Success Metrics

Mål:

  • Find manuals for mindst 50% af tag materialer
  • Reduce scraping tid til under 30 sekunder per produkt
  • Zero false positives (ingen datablade eller sikkerhedsblade)

Status:

  • ✅ Kode opdateret
  • ⏳ Test pending (timeout issues)
  • ⏳ Deployment pending

💡 Observations

Bygma HTML Struktur:

<a ng-click="downloadItem('some-guid', '/files/path/to/Montagevejledning_B7.pdf')">
  Montagevejledning_B7.pdf  <!-- TEXT CONTENT (vigtig!) -->
</a>

Key Insight: PDF filnavnet er ofte i link text, ikke i title attribute!

Alternative Approaches:

  1. Playwright Network Interception - Capture download URLs
  2. JavaScript Evaluation - Execute ng-click og capture PDF URL
  3. Bygma API - Hvis de har en dokumentation API (unlikely)

Konklusion: Scraper er nu mere robust men har performance issues. Næste step er at optimere hastighed og teste med rigtige produkter.