Skip to content

Repository files navigation

My GeekNews Upvotes

한국어 README

A personal search engine for your upvoted articles on GeekNews. This application allows you to scrape your upvoted topics and search through them with a fast, fuzzy search interface that supports Korean Chosung (initial consonant) matching.

Features

  • Incremental Scraping: Efficiently scrapes only new upvoted articles using a Python script.
  • Smart Search: Real-time fuzzy search with Korean Chosung support (e.g., searching "ㄱㄴ" finds "GeekNews").
  • Infinite Scroll: Seamless browsing experience with automatic loading of more articles.
  • Responsive UI: Built with Next.js, Tailwind CSS, and Shadcn UI for a modern, clean look.

Tech Stack

  • Framework: Next.js 15
  • Styling: Tailwind CSS, Shadcn UI
  • Scraping: Python 3 (requests, beautifulsoup4)
  • Utilities: es-hangul (for Korean search logic)

Getting Started

Prerequisites

  • Node.js and npm
  • Python 3
  • A GeekNews account

1. Setup Environment Variables

Copy .env.example to .env and fill in your GeekNews credentials:

cp .env.example .env

Edit .env:

GEEKNEWS_ID=your_id
PASSWORD=your_password
# Optional
# GEEKNEWS_DATA_PATH=/absolute/path/to/your/data.json
# Or use a remote URL (e.g., GitHub Raw, Gist):
# GEEKNEWS_DATA_PATH=https://gist.githubusercontent.com/username/gist_id/raw/geeknews_my_upvotes.json

2. Scrape Upvoted Articles

Install Python dependencies:

pip install requests beautifulsoup4

Run the scraping script:

python3 scrape_geeknews.py

This will create or update data/geeknews_my_upvotes.json with your upvoted articles. New articles will include a date field automatically. Note that the data/ directory is git-ignored to keep your personal data safe.

Backfill Dates for Existing Articles

If you have previously scraped articles without date information, run the backfill script to fetch exact dates from each topic page:

python3 backfill_dates.py

This fetches datePublished from each topic's page metadata. Progress is saved every 50 articles, so you can safely interrupt and resume. Options:

  • --sleep 2 — adjust delay between requests (default: 1 second)
  • --data-path /path/to/data.json — use a custom data file path

Automated Daily Scraping (GitHub Actions)

You can set up a GitHub Actions workflow to scrape automatically once a day:

  1. Go to your repository's Settings → Secrets and variables → Actions
  2. Add two secrets:
    • GEEKNEWS_ID — your GeekNews user ID
    • GEEKNEWS_PASSWORD — your GeekNews password
  3. The workflow (.github/workflows/daily-scrape.yml) runs weekly on Mondays at 09:00 KST (00:00 UTC) and commits updated data to the repository
  4. You can also trigger it manually from the Actions tab → Weekly GeekNews Scrape → Run workflow

3. Run the Application

Install Node.js dependencies:

npm install

Start the development server:

npm run dev

Open http://localhost:3000 with your browser to see the result.

Project Structure

  • scrape_geeknews.py: Python script to scrape upvoted topics.
  • src/app: Next.js app router pages and API routes.
  • src/components: React components (ArticleCard, SearchBar, etc.).
  • src/lib: Utility functions, including the fuzzy search logic.
  • src/services: Data fetching services.

About

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages