Compare commits

3 Commits
v1 ... main

Author SHA1 Message Date
Alex
243bc9b6f5 Add --from flag: travel distance from your location
Some checks are pending
Python application / build (push) Waiting to run
- Geocodes listing locations via Photon (Komoot) API - no rate limits
- Shows distance in km next to location in all output formats
- Uses geodesic distance (geopy) for accurate km calculation
- Caches geocoded locations to avoid repeated API calls
- Works with --top, --compare, --verdict, and default output
- Example: --from "Farum" shows 'København V (20 km)', 'Korsør (94 km)'
- Added geopy to requirements.txt
2026-08-28 16:04:11 +02:00
Copilot
33fdd4dff8 Fix GitHub Actions workflow to handle repositories without tests (#1)
* Initial plan

* Fix workflow to handle missing tests gracefully

Co-authored-by: alexpolo1 <14327609+alexpolo1@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: alexpolo1 <14327609+alexpolo1@users.noreply.github.com>
2026-08-28 15:09:02 +02:00
Alex
14828ddd0f Fix DBA title bug + add 'should I buy it used?' verdict engine
- Fix: title extraction now reads h2 text directly (h2 has no child <a>)
- Fix: link extraction uses a.sf-search-ad-link overlay (was falling through to carousel)
- Fix: skip skeleton/placeholder articles
- Add: --verdict flag with buy_used/buy_new/wait/decent_deal recommendation
- Add: verdict compares median used price vs lowest new price with savings %
- Add: verdict works standalone (no --compare needed)
- Rewrite: README + USAGE docs for the 'go-to' workflow
- Refactor: shared _parse_dba_article() for requests + selenium paths
2026-08-28 15:05:43 +02:00
5 changed files with 466 additions and 207 deletions

View File

@@ -36,4 +36,4 @@ jobs:
flake8 . --count --exit-zero --max-complexity=10 --max-line-length=127 --statistics
- name: Test with pytest
run: |
pytest
pytest || [ $? = 5 ]

110
README.md
View File

@@ -1,56 +1,92 @@
# newolddkprice
DBA vs PriceRunner price comparator
**"Should I buy it used?"** — DBA vs PriceRunner price comparator with verdicts.
This small Python tool searches DBA (dba.dk) and PriceRunner for product
results and prints a side-by-side comparison of titles, prices and links.
Searches [DBA](https://www.dba.dk) for used listings and [PriceRunner](https://www.pricerunner.dk) for new prices, then gives you a side-by-side comparison with a **buy / wait / buy-new** recommendation.
Features
- Fetch DBA search results (requests or Selenium)
- Fetch PriceRunner search results by extracting embedded JSON from the results page
- Filter results by minimum and maximum price
- Compare results in three output formats: plain text table, Markdown, or an ASCII grid
Quick start
Prerequisites
- Python 3.10+ (or 3.8+)
- pip
Install dependencies (recommended in a virtualenv):
## Quick start
```bash
python3 -m venv .venv
source .venv/bin/activate
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
# The go-to command:
python3 dba_pricerunner_scraper.py "playstation 5 konsol" --pricerunner --compare --verdict
```
Run the script (basic example):
Output:
```
==================================================
SHOULD I BUY IT USED? — playstation 5 konsol
==================================================
🤔 DECENT DEAL
Used is only 12% cheaper (~3,500 kr vs 3,995 kr new).
Decent savings but you lose warranty & returns.
savings: [███░░░░░░░░░░░░░░░░░░░░░░░░░░░] 12%
new: 3,995 kr used median: 3,500 kr range: 1984,500 kr
==================================================
```
## Verdicts
| Verdict | When | Meaning |
|---------|------|---------|
| ✅ **BUY USED** | ≥20% savings | Great deal, buy used |
| 🤔 **DECENT DEAL** | 520% savings | Some savings, but you lose warranty |
| 🏷️ **BUY NEW** | <5% savings | Barely any savings new is worth the premium |
| **WAIT** | Used new | Something's off, wait for better deals |
## Options
| Flag | Description |
|------|-------------|
| `--pricerunner` | Also fetch PriceRunner (new) prices |
| `--compare` | Show side-by-side comparison table |
| `--verdict` | Add the "should I buy it used?" recommendation |
| `--format` | `text` (default), `markdown`, or `grid` |
| `--top N` | Show top N results |
| `--min-price` / `--max-price` | Filter by price (e.g. `--min-price 2000`) |
| `--from` | Your location for distance calculation (e.g. `--from Farum`) |
| `--json` | Output raw JSON |
| `--engine selenium` | Use Selenium instead of requests (for JS-heavy pages) |
## Examples
```bash
python3 dba_pricerunner_scraper.py "playstation 5" --pricerunner --compare --format markdown --top 8
# Quick check: is it worth buying used?
python3 dba_pricerunner_scraper.py "iphone 15" --pricerunner --verdict
# Full comparison with markdown table (great for pasting into docs)
python3 dba_pricerunner_scraper.py "samsung galaxy s24" --pricerunner --compare --verdict --format markdown
# Filter out accessories: only look at items 20005000 kr
python3 dba_pricerunner_scraper.py "playstation 5 konsol" --pricerunner --compare --verdict --min-price 2000 --max-price 5000
# Show distances from your location (e.g. Farum)
python3 dba_pricerunner_scraper.py "iphone 15" --from "Farum" --pricerunner --compare --verdict
# ASCII grid for terminal
python3 dba_pricerunner_scraper.py "xbox series x konsol" --pricerunner --compare --verdict --format grid
# Just top 5 DBA listings (no PriceRunner)
python3 dba_pricerunner_scraper.py "playstation 5" --top 5
```
Common options
- `--pricerunner` : fetch PriceRunner results as well
- `--compare` : show a side-by-side comparison (requires `--pricerunner`)
- `--format` : `text` (default), `markdown`, or `grid`
- `--top N` : when used without `--compare`, prints top N DBA results; when used with `--compare` it controls number of rows in the comparison
- `--min-price` / `--max-price` : numeric filters to exclude accessories or outliers
## How it works
Examples
1. **DBA** scrapes search results via `requests` (or Selenium). Extracts title, price, location, and URL from each `<article>` card.
2. **PriceRunner** extracts embedded product JSON from the search results page. Gets the lowest new price for each product.
3. **Verdict** compares the **median used price** (DBA) against the **lowest new price** (PriceRunner) and computes savings %.
- Compare PlayStation 5 results and print a Markdown table:
`python3 dba_pricerunner_scraper.py "playstation 5" --pricerunner --compare --format markdown --top 10`
> **Tip:** Use specific search terms (e.g. "playstation 5 konsol" not "playstation 5") to avoid accessories and games polluting the comparison.
- Compare Google Pixel results and exclude accessories under 3000 DKK:
`python3 dba_pricerunner_scraper.py "google pixel 8" --pricerunner --compare --format grid --min-price 3000 --max-price 12000`
## Requirements
Notes
- PriceRunner pages are parsed by extracting an embedded JSON blob. This works at time of writing but may break if the site changes.
- For very reliable scraping of JS-heavy pages, consider installing Playwright or using the Selenium renderer (the script already contains a Selenium path).
- Python 3.10+
- `requests`, `beautifulsoup4`, `lxml` (see `requirements.txt`)
- Optional: `selenium` + `webdriver-manager` for the `--engine selenium` path
License
This project contains a small utility script; add a license file if you plan to publish.
## License
More detailed usage is in `docs/USAGE.md`.
MIT do whatever you want with it.

View File

@@ -1,11 +1,13 @@
#!/usr/bin/env python3
"""Minimal scraper for dba.dk search results.
"""DBA vs PriceRunner price comparator with "should I buy it used?" verdicts.
Keeps only the DBA scraping code: requests-based fetch and an optional
Selenium renderer. Provides price parsing, simple location heuristics and
CLI for fetching and printing top results.
Searches DBA (dba.dk) for used listings and PriceRunner for new prices,
then gives a side-by-side comparison with a buy/wait recommendation.
Usage: python dba_pricerunner_scraper.py "search terms"
Usage:
python3 dba_pricerunner_scraper.py "playstation 5" --pricerunner --compare --verdict
python3 dba_pricerunner_scraper.py "iphone 15" --pricerunner --compare --format markdown --verdict
python3 dba_pricerunner_scraper.py "playstation 5" --top 10
"""
import sys
@@ -13,10 +15,45 @@ import time
import json
import re
import argparse
import statistics
import requests
from bs4 import BeautifulSoup
from urllib.parse import quote_plus, urljoin
try:
from geopy.distance import geodesic
import requests as _requests
_GEO_CACHE = {}
def _geocode(name):
if not name:
return None
if name in _GEO_CACHE:
return _GEO_CACHE[name]
try:
r = _requests.get(
'https://photon.komoot.io/api/',
params={'q': f'{name}, Danmark'},
headers={'User-Agent': 'dba_pricerunner/1.0'},
timeout=10
)
if r.status_code == 200:
data = r.json()
features = data.get('features', [])
if features:
coords = features[0]['geometry']['coordinates']
result = (coords[1], coords[0]) # [lon, lat] -> (lat, lon)
else:
result = None
else:
result = None
except Exception:
result = None
_GEO_CACHE[name] = result
return result
except ImportError:
def _geocode(name):
return None
try:
from selenium import webdriver
from selenium.webdriver.chrome.service import Service as ChromeService
@@ -34,25 +71,25 @@ class ScraperError(Exception):
pass
# ──────────────────────────────────────────────────────────────
# Price parsing
# ──────────────────────────────────────────────────────────────
def normalize_price(price_str):
"""Extract a numeric price (float) from a string like 'kr. 1.234' or '1.234,00 kr'."""
if not price_str:
return None
# remove currency tokens and non-number suffixes
s = price_str.replace("kr.", "").replace("kr", "").replace("DKK", "")
s = s.replace("\u00A0", " ")
s = s.strip()
# keep only digits, dot, comma and spaces for analysis
import re
cleaned = re.sub(r"[^0-9\.,\s]", "", s)
if not cleaned:
return None
# If there are separators followed by exactly three digits (e.g. '3.250' or '1 234'),
# it's very likely a thousands separator. Remove dots/spaces in that case.
if re.search(r'(?:[\.\s]\d{3})', cleaned):
core = re.sub(r'[\.\s]', '', cleaned)
if re.search(r'(?:[.\s]\d{3})', cleaned):
core = re.sub(r'[.\s]', '', cleaned)
core = core.replace(',', '.')
core = ''.join(ch for ch in core if (ch.isdigit() or ch == '.'))
try:
@@ -67,16 +104,13 @@ def normalize_price(price_str):
if last_dot == -1 and last_comma == -1:
decimal_sep = None
elif last_dot > last_comma:
# dot occurs later
if len(cleaned) - last_dot - 1 in (1, 2, 3):
decimal_sep = '.'
else:
if len(cleaned) - last_comma - 1 in (1, 2, 3):
decimal_sep = ','
# remove thousands separators (either '.' or ',') except the decimal separator
if decimal_sep is None:
# just remove spaces and separators
digits = re.sub(r"[\s\.,]", "", cleaned)
try:
return float(digits)
@@ -84,11 +118,9 @@ def normalize_price(price_str):
return None
else:
if decimal_sep == '.':
# remove commas and spaces, keep dot
core = re.sub(r"[\s,]", "", cleaned)
else:
# decimal_sep == ',' -> remove dots and spaces, replace comma with dot
core = re.sub(r"[\.\s]", "", cleaned)
core = re.sub(r"[.\s]", "", cleaned)
core = core.replace(',', '.')
core = ''.join(ch for ch in core if (ch.isdigit() or ch == '.'))
try:
@@ -98,24 +130,17 @@ def normalize_price(price_str):
def extract_price_string(text):
"""Return a short price string from a larger text blob, e.g. '3.999 kr.' or '150 kr.'
Uses a regex to find a Danish-style price (thousands sep '.' or space, decimal comma).
"""
"""Return a short price string from a larger text blob, e.g. '3.999 kr.' or '150 kr.'"""
if not text:
return ''
# look for number patterns optionally followed/preceded by currency
import re
m = re.search(r"(\d{1,3}(?:[\.\s]\d{3})*(?:,\d{1,2})?)\s*(kr\.?|DKK)?", text, flags=re.IGNORECASE)
m = re.search(r"(\d{1,3}(?:[.\s]\d{3})*(?:,\d{1,2})?)\s*(kr\.?|DKK)?", text, flags=re.IGNORECASE)
if not m:
# fallback: try a simpler digit sequence
m = re.search(r"(\d+[\d\.\s,]*)", text)
m = re.search(r"(\d+[\d.\s,]*)", text)
if not m:
return ''
num = m.group(1)
cur = m.group(2) or 'kr.'
cur = cur.strip()
# normalize currency display
if cur.upper() == 'DKK':
cur = 'DKK'
elif not cur:
@@ -123,23 +148,75 @@ def extract_price_string(text):
return f"{num} {cur}".strip()
def extract_location_from_element(el):
"""Heuristic: try to find a location string near a listing element."""
if el is None:
return ''
for cls in ('cAdList__location', 'ad-location', 'dba-location', 'location', 'by', 'region'):
node = el.find(class_=cls)
if node and node.get_text(strip=True):
return node.get_text(strip=True)
for small in el.find_all(['small', 'span', 'p']):
txt = small.get_text(strip=True)
if txt and any(ch.isdigit() for ch in txt) is False and len(txt) < 60:
return txt
return ''
# ──────────────────────────────────────────────────────────────
# DBA scraping
# ──────────────────────────────────────────────────────────────
def _parse_dba_article(el, origin_coords=None):
"""Extract title, price, url, location from a DBA <article> element.
If origin_coords is provided, also computes travel distance from origin.
"""
# Title: h2 text directly (the h2 has NO child <a> in current markup)
h2 = el.select_one('h2')
title = h2.get_text(strip=True) if h2 else ''
# Link: the sf-search-ad-link overlay (direct child of <article>)
link_el = el.select_one('a.sf-search-ad-link')
link = urljoin('https://www.dba.dk', link_el['href']) if link_el and link_el.get('href') else None
# Price: first span inside the flex justify-between div
price = ''
price_div = el.select_one('div.flex.justify-between')
if price_div:
price = price_div.get_text(strip=True)
if not price:
# fallback: scan all spans
for tag in el.find_all(['span', 'div', 'p']):
txt = tag.get_text(' ', strip=True)
if not txt:
continue
pstr = extract_price_string(txt)
if pstr:
price = pstr
break
# Location: span with whitespace-nowrap truncate mr-8
location = ''
loc_el = el.select_one('span.whitespace-nowrap.truncate')
if loc_el and loc_el.get_text(strip=True):
location = loc_el.get_text(strip=True)
else:
loc_block = el.select_one('.text-xs.s-text-subtle')
if loc_block:
sp = loc_block.select_one('span')
if sp and sp.get_text(strip=True):
location = sp.get_text(strip=True)
# Distance from origin (if provided)
distance_km = None
if origin_coords and location:
coords = _geocode(location)
if coords:
distance_km = geodesic(origin_coords, coords).km
price_num = normalize_price(price)
return {
"site": "dba",
"title": title,
"price": price,
"price_num": price_num,
"url": link,
"location": location,
"distance_km": distance_km,
}
def search_dba_requests(query, max_results=10):
def search_dba_requests(query, max_results=10, origin=None):
"""Search dba.dk (requests) and return list of dicts with title, price, url, location."""
origin_coords = None
if origin:
origin_coords = _geocode(origin)
q = quote_plus(query)
url = f"https://www.dba.dk/recommerce/forsale/search?q={q}"
r = requests.get(url, headers=HEADERS, timeout=15)
@@ -149,43 +226,23 @@ def search_dba_requests(query, max_results=10):
items = []
results = soup.select('article') or soup.select('.cAdList__item') or soup.select('.dba-result')
for el in results[:max_results]:
a = el.select_one('h2 a') or el.select_one('a.sf-search-ad-link') or el.select_one('a')
title = a.get_text(strip=True) if a else ''
link = urljoin('https://www.dba.dk', a['href']) if a and a.get('href') else None
price = ''
for tag in el.find_all(['span', 'div', 'p']):
txt = tag.get_text(' ', strip=True)
if not txt:
continue
pstr = extract_price_string(txt)
if pstr:
price = pstr
break
location = ''
loc_block = el.select_one('.text-xs.s-text-subtle') or el.select_one('.cAdList__location')
if loc_block:
sp = loc_block.select_one('span')
if sp and sp.get_text(strip=True):
location = sp.get_text(strip=True)
if not location:
location = extract_location_from_element(el) or ''
price_num = normalize_price(price)
items.append({
"site": "dba",
"title": title,
"price": price,
"price_num": price_num,
"url": link,
"location": location,
})
# skip skeleton/placeholder articles (no text content)
if not el.get_text(strip=True):
continue
item = _parse_dba_article(el, origin_coords)
if item['title']:
items.append(item)
time.sleep(1)
return items
def search_dba_selenium(query, max_results=10, headless=True):
def search_dba_selenium(query, max_results=10, headless=True, origin=None):
"""Render DBA with Selenium and extract the same fields as requests path."""
if webdriver is None:
raise ScraperError('Selenium or webdriver-manager not installed')
origin_coords = None
if origin:
origin_coords = _geocode(origin)
opts = ChromeOptions()
if headless:
opts.add_argument('--headless=new')
@@ -201,28 +258,11 @@ def search_dba_selenium(query, max_results=10, headless=True):
items = []
results = soup.select('article') or soup.select('.cAdList__item') or soup.select('.dba-result')
for el in results[:max_results]:
a = el.select_one('h2 a') or el.select_one('a.sf-search-ad-link') or el.select_one('a')
title = a.get_text(strip=True) if a else ''
link = urljoin('https://www.dba.dk', a['href']) if a and a.get('href') else None
price = ''
for tag in el.find_all(['span', 'div', 'p']):
txt = tag.get_text(' ', strip=True)
if not txt:
continue
pstr = extract_price_string(txt)
if pstr:
price = pstr
break
location = ''
loc_block = el.select_one('.text-xs.s-text-subtle') or el.select_one('.cAdList__location')
if loc_block:
sp = loc_block.select_one('span')
if sp and sp.get_text(strip=True):
location = sp.get_text(strip=True)
if not location:
location = extract_location_from_element(el) or ''
price_num = normalize_price(price)
items.append({'site': 'dba', 'title': title, 'price': price, 'price_num': price_num, 'url': link, 'location': location})
if not el.get_text(strip=True):
continue
item = _parse_dba_article(el, origin_coords)
if item['title']:
items.append(item)
time.sleep(1)
return items
finally:
@@ -232,19 +272,12 @@ def search_dba_selenium(query, max_results=10, headless=True):
pass
def sort_items_by_price(items):
def keyfn(it):
v = it.get('price_num')
return v if v is not None else float('inf')
return sorted(items, key=keyfn)
# ──────────────────────────────────────────────────────────────
# PriceRunner scraping
# ──────────────────────────────────────────────────────────────
def _extract_json_array(text, key):
"""Find a JSON array by key in a large HTML/JS blob and return the array text.
Scans for '"key":[' and returns the bracketed array (handles nested brackets
and strings). Returns None on failure.
"""
"""Find a JSON array by key in a large HTML/JS blob and return the array text."""
needle = f'"{key}":['
idx = text.find(needle)
if idx == -1:
@@ -273,10 +306,7 @@ def _extract_json_array(text, key):
def search_pricerunner_requests(query, max_results=10):
"""Fetch PriceRunner search page and extract embedded product JSON (requests).
Returns list of dicts: title, price, price_num, url, site='pricerunner'.
"""
"""Fetch PriceRunner search page and extract embedded product JSON (requests)."""
q = quote_plus(query)
url = f"https://www.pricerunner.dk/results?q={q}"
r = requests.get(url, headers=HEADERS, timeout=15)
@@ -285,7 +315,6 @@ def search_pricerunner_requests(query, max_results=10):
arr_text = _extract_json_array(r.text, 'products')
items = []
if not arr_text:
# fallback: try to find simple product blocks
soup = BeautifulSoup(r.text, 'lxml')
cards = soup.select('.product, .product-item, .search-result')
for el in cards[:max_results]:
@@ -308,7 +337,6 @@ def search_pricerunner_requests(query, max_results=10):
lp = p.get('lowestPrice') or {}
if isinstance(lp, dict):
price = lp.get('amount')
# price may be a string like '3289.00' or None
price_str = f"{price} {lp.get('currency','')}".strip() if price else ''
path = p.get('url')
full = urljoin('https://www.pricerunner.dk', path) if path else None
@@ -318,19 +346,138 @@ def search_pricerunner_requests(query, max_results=10):
return items
def print_top_results(items, n=5):
print(f"Top {n} DBA results (title — price — location):\n")
for it in items[:n]:
loc = it.get('location') or ''
print(f"- {it.get('title','')}")
print(f" price: {it.get('price','')}\n location: {loc}\n url: {it.get('url','')}\n")
# ──────────────────────────────────────────────────────────────
# Verdict engine — "should I buy it used?"
# ──────────────────────────────────────────────────────────────
def compute_verdict(dba_items, pr_items):
"""Compare used (DBA) vs new (PriceRunner) prices and produce a recommendation.
Returns a dict with:
new_price, used_price, used_median, savings, savings_pct,
verdict (one of: 'buy_used', 'buy_new', 'wait', 'mixed'),
reason (human-readable explanation)
"""
new_prices = [it['price_num'] for it in pr_items if it.get('price_num') is not None]
used_prices = [it['price_num'] for it in dba_items if it.get('price_num') is not None]
if not new_prices or not used_prices:
return {
'new_price': None, 'used_price': None, 'used_median': None,
'savings': None, 'savings_pct': None,
'verdict': 'no_data',
'reason': 'Not enough price data to compare. Try a more specific search term.'
}
new_price = min(new_prices)
used_median = statistics.median(used_prices)
used_low = min(used_prices)
used_high = max(used_prices)
savings = new_price - used_median
savings_pct = (savings / new_price * 100) if new_price > 0 else 0
# Decide the verdict
if savings_pct >= 40:
verdict = 'buy_used'
reason = (f"Used is {savings_pct:.0f}% cheaper than new "
f"(~{used_median:,.0f} kr vs {new_price:,.0f} kr new). "
f"Great deal — buy used!")
elif savings_pct >= 20:
verdict = 'buy_used'
reason = (f"Used saves you {savings_pct:.0f}% "
f"(~{used_median:,.0f} kr vs {new_price:,.0f} kr new). "
f"Solid savings, worth considering used.")
elif savings_pct >= 5:
verdict = 'mixed'
reason = (f"Used is only {savings_pct:.0f}% cheaper "
f"(~{used_median:,.0f} kr vs {new_price:,.0f} kr new). "
f"Decent savings but you lose warranty & returns.")
elif savings_pct > 0:
verdict = 'buy_new'
reason = (f"Used barely saves anything ({savings_pct:.0f}%) "
f"(~{used_median:,.0f} kr vs {new_price:,.0f} kr new). "
f"New comes with warranty — probably worth the small premium.")
else:
verdict = 'wait'
reason = (f"Used prices ({used_low:,.0f}{used_high:,.0f} kr) are at or above "
f"new price ({new_price:,.0f} kr). Something's off — wait for better deals.")
return {
'new_price': new_price,
'used_price': used_median,
'used_median': used_median,
'used_low': used_low,
'used_high': used_high,
'savings': savings,
'savings_pct': savings_pct,
'verdict': verdict,
'reason': reason,
}
def format_verdict(v, fmt='text'):
"""Format the verdict dict for display."""
if v['verdict'] == 'no_data':
return v['reason']
icon = {
'buy_used': '✅ BUY USED',
'buy_new': '🏷️ BUY NEW',
'mixed': '🤔 DECENT DEAL',
'wait': '⏳ WAIT',
}.get(v['verdict'], '')
if fmt == 'markdown':
lines = [
f"**{icon}** — {v['reason']}",
f"| New (PriceRunner) | Used (DBA median) | You save |",
f"|---|---|---|",
f"| {v['new_price']:,.0f} kr | {v['used_median']:,.0f} kr | {v['savings']:,.0f} kr ({v['savings_pct']:.0f}%) |",
]
return '\n'.join(lines)
# text / grid
bar_len = 30
ratio = min(v['savings_pct'] / 100, 1.0) if v['savings_pct'] > 0 else 0
bar = '' * int(ratio * bar_len) + '' * (bar_len - int(ratio * bar_len))
lines = [
f" {icon}",
f" {v['reason']}",
f" savings: [{bar}] {v['savings_pct']:.0f}%",
f" new: {v['new_price']:,.0f} kr used median: {v['used_median']:,.0f} kr "
f"range: {v['used_low']:,.0f}{v['used_high']:,.0f} kr",
]
return '\n'.join(lines)
# ──────────────────────────────────────────────────────────────
# Output formatting
# ──────────────────────────────────────────────────────────────
def sort_items_by_price(items):
def keyfn(it):
v = it.get('price_num')
return v if v is not None else float('inf')
return sorted(items, key=keyfn)
def _short(s, n=60):
if not s:
return ''
s = ' '.join(s.split())
return s if len(s) <= n else s[: n-3] + '...'
return s if len(s) <= n else s[:n-3] + '...'
def print_top_results(items, n=5):
print(f"Top {n} DBA results (title — price — location):\n")
for it in items[:n]:
loc = it.get('location') or ''
dist = it.get('distance_km')
if dist is not None:
loc = f"{loc} ({dist:.0f} km)"
print(f"- {it.get('title','')}")
print(f" price: {it.get('price','')}\n location: {loc}\n url: {it.get('url','')}\n")
def print_comparison_table(dba_items, pr_items, n=5):
@@ -340,22 +487,23 @@ def print_comparison_table(dba_items, pr_items, n=5):
for i in range(maxrows):
left = dba_items[i] if i < len(dba_items) else None
right = pr_items[i] if i < len(pr_items) else None
loc = left.get('location','') if left else ''
dist = left.get('distance_km') if left else None
if dist is not None:
loc = f"{loc} ({dist:.0f} km)"
rows.append((
str(i+1),
_short(left.get('title','')) if left else '',
left.get('price','') if left else '',
left.get('location','') if left else '',
loc,
_short(right.get('title','')) if right else '',
right.get('price','') if right else '',
))
# column widths
widths = [3, 50, 12, 12, 50, 12]
widths = [3, 50, 12, 18, 50, 12]
hdr = ('#', 'DBA title', 'DBA price', 'DBA loc', 'PriceRunner title', 'PR price')
sep = ' | '
def fmt(row):
return sep.join(row[i].ljust(widths[i]) for i in range(len(row)))
print('\nComparison table (DBA vs PriceRunner):')
print(fmt(hdr))
print('-' * (sum(widths) + len(sep) * (len(widths)-1)))
@@ -365,12 +513,9 @@ def print_comparison_table(dba_items, pr_items, n=5):
def print_comparison_markdown(dba_items, pr_items, n=10):
"""Print a markdown table with two columns: DBA and PriceRunner.
Each cell contains a linked title (if URL available), price and optional location.
"""
"""Print a markdown table with two columns: DBA and PriceRunner."""
lines = []
lines.append("| DBA | PriceRunner |")
lines.append("| DBA (used) | PriceRunner (new) |")
lines.append("|-----|------------|")
def cell(it):
@@ -384,6 +529,9 @@ def print_comparison_markdown(dba_items, pr_items, n=10):
title_md = title
price = it.get('price','')
loc = it.get('location','')
dist = it.get('distance_km')
if dist is not None:
loc = f"{loc} ({dist:.0f} km)"
parts = [title_md, price]
if loc:
parts.append(loc)
@@ -411,8 +559,10 @@ def print_comparison_grid(dba_items, pr_items, n=10):
title = it.get('title','')
price = it.get('price','')
loc = it.get('location','')
dist = it.get('distance_km')
if dist is not None:
loc = f"{loc} ({dist:.0f} km)"
url = it.get('url','') or ''
# shorten URL for display
if url and len(url) > 80:
url = url[:77] + '...'
lines = [title, price]
@@ -420,15 +570,12 @@ def print_comparison_grid(dba_items, pr_items, n=10):
lines.append(loc)
if url:
lines.append(url)
# wrap lines to preferred width later
return lines
left_col.append(cell_lines(l))
right_col.append(cell_lines(r))
# preferred maximum widths
LEFT_MAX = 60
RIGHT_MAX = 80
# wrap each cell's lines to the column max width
def wrap_lines(block, width):
import textwrap
wrapped = []
@@ -436,7 +583,6 @@ def print_comparison_grid(dba_items, pr_items, n=10):
if not line:
wrapped.append('')
else:
# use textwrap to preserve words
for w in textwrap.wrap(line, width=width) or ['']:
wrapped.append(w)
return wrapped
@@ -444,7 +590,6 @@ def print_comparison_grid(dba_items, pr_items, n=10):
left_col = [wrap_lines(b, LEFT_MAX) for b in left_col]
right_col = [wrap_lines(b, RIGHT_MAX) for b in right_col]
# compute column widths (use the max of wrapped lines but cap at the MAX)
left_w = min(max((len(line) for block in left_col for line in block), default=10), LEFT_MAX)
right_w = min(max((len(line) for block in right_col for line in block), default=10), RIGHT_MAX)
sep = ' | '
@@ -462,44 +607,62 @@ def print_comparison_grid(dba_items, pr_items, n=10):
print()
# ──────────────────────────────────────────────────────────────
# Main
# ──────────────────────────────────────────────────────────────
def main():
parser = argparse.ArgumentParser(description='Fetch DBA search results')
parser = argparse.ArgumentParser(
description='DBA vs PriceRunner price comparator — "should I buy it used?"'
)
parser.add_argument('query', nargs='+')
parser.add_argument('--engine', choices=['requests', 'selenium'], default='requests')
parser.add_argument('--max', type=int, default=15)
parser.add_argument('--json', action='store_true')
parser.add_argument('--top', type=int, default=0, help='Print top N DBA results with price and location')
parser.add_argument('--pricerunner', action='store_true', help='Fetch PriceRunner search results as well')
parser.add_argument('--compare', action='store_true', help='Print a comparison table of DBA vs PriceRunner (uses --pricerunner)')
parser.add_argument('--format', choices=['text', 'markdown', 'grid'], default='text', help='Output format for comparison table')
parser.add_argument('--min-price', type=str, default=None, help='Filter out items below this price (e.g. 500 or "3.000")')
parser.add_argument('--max-price', type=str, default=None, help='Filter out items above this price')
parser.add_argument('--top', type=int, default=0,
help='Print top N DBA results with price and location')
parser.add_argument('--pricerunner', action='store_true',
help='Fetch PriceRunner search results as well')
parser.add_argument('--compare', action='store_true',
help='Print a comparison table of DBA vs PriceRunner (uses --pricerunner)')
parser.add_argument('--format', choices=['text', 'markdown', 'grid'], default='text',
help='Output format for comparison table')
parser.add_argument('--verdict', action='store_true',
help='Add a "should I buy it used?" recommendation')
parser.add_argument('--min-price', type=str, default=None,
help='Filter out items below this price (e.g. 500 or "3.000")')
parser.add_argument('--max-price', type=str, default=None,
help='Filter out items above this price')
parser.add_argument('--from', dest='origin', type=str, default=None,
help='Your location for distance calculation (e.g. "Farum")')
args = parser.parse_args()
query = ' '.join(args.query)
print(f"Searching for: {query}\n")
# ── Fetch DBA ──
try:
if args.engine == 'requests':
dba = search_dba_requests(query, max_results=args.max)
dba = search_dba_requests(query, max_results=args.max, origin=args.origin)
else:
dba = search_dba_selenium(query, max_results=args.max)
dba = search_dba_selenium(query, max_results=args.max, origin=args.origin)
except ScraperError as e:
print("Error while scraping DBA:", e)
sys.exit(1)
# ── Fetch PriceRunner ──
pr = []
if args.pricerunner or args.compare:
if args.pricerunner or args.compare or args.verdict:
try:
pr = search_pricerunner_requests(query, max_results=args.max)
except ScraperError as e:
print('Error while scraping PriceRunner:', e)
# parse min/max price args into floats using normalize_price
# ── Price filters ──
def _parse_price_arg(s):
if s is None:
return None
v = normalize_price(s)
if v is None:
# try to strip currency and commas
try:
s2 = s.replace('.', '').replace(',', '.')
return float(''.join(ch for ch in s2 if (ch.isdigit() or ch == '.')))
@@ -527,36 +690,60 @@ def main():
print(f"Applied price filter: min={minp} max={maxp}. DBA: {before_d}->{len(dba)}, PR: {before_p}->{len(pr)}\n")
print(f"Found {len(dba)} items on DBA\n")
# ── JSON output ──
if args.json:
print(json.dumps(dba, ensure_ascii=False, indent=2))
return
# If the user asked for a simple top-N DBA listing (without compare), show and exit.
# ── Top-N DBA only ──
if args.top and args.top > 0 and not args.compare:
print_top_results(dba, n=args.top)
return
if args.pricerunner and not args.compare:
# ── PriceRunner only (no compare) ──
if args.pricerunner and not args.compare and not args.verdict:
print(f"Found {len(pr)} items on PriceRunner\n")
if args.json:
print(json.dumps(pr, ensure_ascii=False, indent=2))
return
# ── Verdict ──
if args.verdict:
v = compute_verdict(dba, pr)
print(f"\n{'='*50}")
print(f" SHOULD I BUY IT USED? — {query}")
print(f"{'='*50}")
print(format_verdict(v, fmt=args.format))
print(f"{'='*50}\n")
# ── Comparison table ──
if args.compare:
# sort both lists by price_num for a reasonable alignment
d_sorted = sort_items_by_price(dba)
p_sorted = sort_items_by_price(pr)
n = args.top if args.top and args.top > 0 else 10
if args.format == 'markdown':
print_comparison_markdown(d_sorted, p_sorted, n=args.top if args.top and args.top>0 else 10)
print_comparison_markdown(d_sorted, p_sorted, n=n)
elif args.format == 'grid':
print_comparison_grid(d_sorted, p_sorted, n=args.top if args.top and args.top>0 else 10)
print_comparison_grid(d_sorted, p_sorted, n=n)
else:
print_comparison_table(d_sorted, p_sorted, n=args.top if args.top and args.top>0 else 10)
print_comparison_table(d_sorted, p_sorted, n=n)
return
# ── Top-N DBA only (no compare, no verdict) ──
if args.top and args.top > 0:
print_top_results(dba, n=args.top)
return
# ── Default: sorted DBA list ──
dba_sorted = sort_items_by_price(dba)
print(f"DBA — top {min(len(dba_sorted), args.max)} by price:")
for it in dba_sorted[:args.max]:
print(f"- {it.get('title','')}{it.get('price','')}{it.get('location','')}{it.get('url','')}")
loc = it.get('location','')
dist = it.get('distance_km')
if dist is not None:
loc = f"{loc} ({dist:.0f} km)"
print(f"- {it.get('title','')}{it.get('price','')}{loc}{it.get('url','')}")
if __name__ == '__main__':

View File

@@ -1,35 +1,70 @@
# Usage
This document expands on how to run `dba_pricerunner_scraper.py`.
Basic invocation
## The go-to command
```bash
python3 dba_pricerunner_scraper.py "search terms" --pricerunner --compare --format grid --top 8
python3 dba_pricerunner_scraper.py "YOUR SEARCH" --pricerunner --compare --verdict
```
Options
This fetches DBA (used) + PriceRunner (new) results, prints a comparison table, and gives you a verdict:
- `--engine`: `requests` (default) or `selenium`.
- `--pricerunner`: fetch PriceRunner results.
- `--compare`: show side-by-side comparison.
- `--format`: `text` | `markdown` | `grid`.
- `--min-price` / `--max-price`: filter by price.
- **BUY USED** — used is 20%+ cheaper
- 🤔 **DECENT DEAL** — used is 520% cheaper
- 🏷️ **BUY NEW** — used barely saves anything
- **WAIT** — used is at or above new price
Examples
## All options
- Top 5 DBA results only:
| Flag | Default | Description |
|------|---------|-------------|
| `query` | (required) | Search terms (quote multi-word) |
| `--engine` | `requests` | `requests` or `selenium` |
| `--max` | `15` | Max results to fetch |
| `--top N` | — | Show top N results |
| `--pricerunner` | off | Also fetch PriceRunner prices |
| `--compare` | off | Show side-by-side comparison |
| `--verdict` | off | Add buy/wait recommendation |
| `--format` | `text` | `text`, `markdown`, or `grid` |
| `--min-price` | — | Filter: exclude items below this price |
| `--max-price` | — | Filter: exclude items above this price |
| `--json` | off | Output raw JSON instead of tables |
## Examples
```bash
# Quick verdict (no table)
python3 dba_pricerunner_scraper.py "iphone 15" --pricerunner --verdict
# Full comparison, markdown format
python3 dba_pricerunner_scraper.py "samsung galaxy s24" --pricerunner --compare --verdict --format markdown
# Filter out cheap accessories
python3 dba_pricerunner_scraper.py "playstation 5 konsol" --pricerunner --compare --verdict --min-price 2000
# ASCII grid for terminal
python3 dba_pricerunner_scraper.py "xbox series x konsol" --pricerunner --compare --verdict --format grid
# Top 5 DBA only (no PriceRunner)
python3 dba_pricerunner_scraper.py "playstation 5" --top 5
# JSON output for scripting
python3 dba_pricerunner_scraper.py "iphone 15" --json
```
- Compare results and save markdown output:
## Tips
- **Be specific with search terms.** "playstation 5 konsol" gives much better results than "playstation 5" (which picks up games, controllers, and cases).
- **Use `--min-price` to filter accessories.** E.g. `--min-price 1000` for phones, `--min-price 2000` for consoles.
- **The verdict uses the median used price**, not the cheapest. This avoids outliers (e.g. one broken console at 500 kr dragging the median down).
- **PriceRunner picks the lowest new price** across all sellers. This is your "best case" new price.
## Selenium engine
If DBA starts blocking `requests`, try:
```bash
python3 dba_pricerunner_scraper.py "playstation 5" --pricerunner --compare --format markdown --top 10 > results.md
pip install selenium webdriver-manager
python3 dba_pricerunner_scraper.py "playstation 5" --engine selenium --pricerunner --compare --verdict
```
Notes on Selenium
If you pick `--engine selenium`, install a compatible browser (Chrome) and ensure `chromedriver` is available. `webdriver-manager` can install the driver automatically but still requires a Chrome binary.
This renders the page in a headless Chrome browser. You need Chrome installed on the system.

View File

@@ -1,5 +1,6 @@
requests
beautifulsoup4
lxml
geopy
webdriver-manager
selenium