A personal search engine for your upvoted articles on GeekNews. This application allows you to scrape your upvoted topics and search through them with a fast, fuzzy search interface that supports Korean Chosung (initial consonant) matching.
- Incremental Scraping: Efficiently scrapes only new upvoted articles using a Python script.
- Smart Search: Real-time fuzzy search with Korean Chosung support (e.g., searching "ㄱㄴ" finds "GeekNews").
- Infinite Scroll: Seamless browsing experience with automatic loading of more articles.
- Responsive UI: Built with Next.js, Tailwind CSS, and Shadcn UI for a modern, clean look.
- Framework: Next.js 15
- Styling: Tailwind CSS, Shadcn UI
- Scraping: Python 3 (
requests,beautifulsoup4) - Utilities:
es-hangul(for Korean search logic)
- Node.js and npm
- Python 3
- A GeekNews account
Copy .env.example to .env and fill in your GeekNews credentials:
cp .env.example .envEdit .env:
GEEKNEWS_ID=your_id
PASSWORD=your_password
# Optional
# GEEKNEWS_DATA_PATH=/absolute/path/to/your/data.json
# Or use a remote URL (e.g., GitHub Raw, Gist):
# GEEKNEWS_DATA_PATH=https://gist.githubusercontent.com/username/gist_id/raw/geeknews_my_upvotes.jsonInstall Python dependencies:
pip install requests beautifulsoup4Run the scraping script:
python3 scrape_geeknews.pyThis will create or update data/geeknews_my_upvotes.json with your upvoted articles. New articles will include a date field automatically. Note that the data/ directory is git-ignored to keep your personal data safe.
If you have previously scraped articles without date information, run the backfill script to fetch exact dates from each topic page:
python3 backfill_dates.pyThis fetches datePublished from each topic's page metadata. Progress is saved every 50 articles, so you can safely interrupt and resume. Options:
--sleep 2— adjust delay between requests (default: 1 second)--data-path /path/to/data.json— use a custom data file path
You can set up a GitHub Actions workflow to scrape automatically once a day:
- Go to your repository's Settings → Secrets and variables → Actions
- Add two secrets:
GEEKNEWS_ID— your GeekNews user IDGEEKNEWS_PASSWORD— your GeekNews password
- The workflow (
.github/workflows/daily-scrape.yml) runs weekly on Mondays at 09:00 KST (00:00 UTC) and commits updated data to the repository - You can also trigger it manually from the Actions tab → Weekly GeekNews Scrape → Run workflow
Install Node.js dependencies:
npm installStart the development server:
npm run devOpen http://localhost:3000 with your browser to see the result.
scrape_geeknews.py: Python script to scrape upvoted topics.src/app: Next.js app router pages and API routes.src/components: React components (ArticleCard, SearchBar, etc.).src/lib: Utility functions, including the fuzzy search logic.src/services: Data fetching services.