A sophisticated, CLI-based visualization tool for researchers who need their screens to look busy.
"Efficiency is doing things right; effectiveness is doing the right things."
— Peter Drucker (and every researcher waiting for a model to converge)
In the era of Generative AI, training times for Large Language Models (e.g., Llama-3, Mistral) have increased exponentially. Researchers often need to manually verify "text-based data streams" to ensure data integrity and model alignment during the fine-tuning process.
DTM provides a robust interface to "visualize" your local datasets (in .epub format) disguised as standard HuggingFace/PyTorch training logs. This ensures your terminal activity remains indistinguishable from legitimate research work, keeping your screen busy and professional.
-
LLM Training Simulation: Automatically generates realistic metrics including
Perplexity (PPL),GradNorm,Tokens/s, and decreasingLossto mimic a healthy convergence trajectory. -
Smart Resume System: Automatically saves your "training state" (global step, epoch) to
.llm_training_state.json. Resume exactly where you left off after a crash or manual interruption. -
Interactive Controls:
- Stepping: Process data chunk-by-chunk.
- Boss Key (
err): Instantly triggers a fakeRuntimeError: CUDA out of memorystack trace (simulating a crash inside a Transformer forward pass), perfect for explaining sudden workflow interruptions.
-
Universal Dataset Support: Reads standard
.epubfiles via a robust fallback parser that handles complex chapter structures.
# Clone the repository git clone https://github.com/your-username/DeepTrain-Monitor.git cd DeepTrain-Monitor # Install dependencies pip install beautifulsoup4
To avoid copyright issues and ensure safety, this repository does not include copyrighted books.
You should place your own .epub file (e.g., Moby Dick, technical docs, or your "target data") into the data/ directory.
Recommended: Use public domain books from Project Gutenberg for testing.
You can run the script with the default path or specify your own dataset via command line arguments.
Default Mode (Looks for ./data/default_novel.epub):
python train_llm.py
Custom Dataset Mode:
python train_llm.py --path "./data/my_research_material.epub"
| Command | Description | Scientific Explanation |
|---|---|---|
Enter |
Next Chunk | Process the next sequence of tokens. |
q |
Quit | Graceful interruption (saves checkpoint). |
err |
Panic Mode | Simulate a fatal CUDA OOM exception (Boss Key). |
Your terminal will display logs indistinguishable from a standard Distributed Data Parallel (DDP) training loop:
[INFO] Initializing Distributed Data Parallel (DDP)... [INFO] Loading Tokenizer: Llama-3-8B-Instruct... [INFO] Model loaded on cuda:0. Trainable params: 6,738,415,616 ---------------------------------------------------------------------------------------------------- Step | Epoch | Metrics | Log Message ---------------------------------------------------------------------------------------------------- [1250 ] [1 ] Loss: 1.8420 | PPL: 6.31 | GradNorm: 1.12 | 2950 tok/s | msg: Call me Ishmael. [1251 ] [1 ] Loss: 1.8385 | PPL: 6.29 | GradNorm: 0.98 | 3012 tok/s | msg: Some years ago...
Pull requests are welcome. For major changes, please open an issue first to discuss what you would like to change.
This tool is for educational and research purposes only.
- Copyright: Please ensure you have the right to read/use the
.epubfiles you load. The author does not endorse piracy. - Usage: The author is not responsible for any extended graduation dates, missed deadlines, or confusion caused by "exceptionally fast reading speeds" during working hours.
Happy "Training"! 🎓