Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

DeepTrain-Monitor (DTM) 🚀

Python PyTorch License Build

A sophisticated, CLI-based visualization tool for researchers who need their screens to look busy.


"Efficiency is doing things right; effectiveness is doing the right things."
— Peter Drucker (and every researcher waiting for a model to converge)

📖 Abstract

In the era of Generative AI, training times for Large Language Models (e.g., Llama-3, Mistral) have increased exponentially. Researchers often need to manually verify "text-based data streams" to ensure data integrity and model alignment during the fine-tuning process.

DTM provides a robust interface to "visualize" your local datasets (in .epub format) disguised as standard HuggingFace/PyTorch training logs. This ensures your terminal activity remains indistinguishable from legitimate research work, keeping your screen busy and professional.


✨ Key Features

  • LLM Training Simulation: Automatically generates realistic metrics including Perplexity (PPL), GradNorm, Tokens/s, and decreasing Loss to mimic a healthy convergence trajectory.
  • Smart Resume System: Automatically saves your "training state" (global step, epoch) to .llm_training_state.json. Resume exactly where you left off after a crash or manual interruption.
  • Interactive Controls:
    • Stepping: Process data chunk-by-chunk.
    • Boss Key (err): Instantly triggers a fake RuntimeError: CUDA out of memory stack trace (simulating a crash inside a Transformer forward pass), perfect for explaining sudden workflow interruptions.
  • Universal Dataset Support: Reads standard .epub files via a robust fallback parser that handles complex chapter structures.

🛠 Installation

# Clone the repository
git clone https://github.com/your-username/DeepTrain-Monitor.git
cd DeepTrain-Monitor

# Install dependencies
pip install beautifulsoup4

🚀 Usage

1. Dataset Preparation

To avoid copyright issues and ensure safety, this repository does not include copyrighted books. You should place your own .epub file (e.g., Moby Dick, technical docs, or your "target data") into the data/ directory.

Recommended: Use public domain books from Project Gutenberg for testing.

2. Run the Monitor

You can run the script with the default path or specify your own dataset via command line arguments.

Default Mode (Looks for ./data/default_novel.epub):

python train_llm.py

Custom Dataset Mode:

python train_llm.py --path "./data/my_research_material.epub"

🎮 Controls (Interactive Mode)

Command Description Scientific Explanation
Enter Next Chunk Process the next sequence of tokens.
q Quit Graceful interruption (saves checkpoint).
err Panic Mode Simulate a fatal CUDA OOM exception (Boss Key).

📊 Output Example

Your terminal will display logs indistinguishable from a standard Distributed Data Parallel (DDP) training loop:

[INFO] Initializing Distributed Data Parallel (DDP)...
[INFO] Loading Tokenizer: Llama-3-8B-Instruct...
[INFO] Model loaded on cuda:0. Trainable params: 6,738,415,616
----------------------------------------------------------------------------------------------------
Step     | Epoch  | Metrics                                                 | Log Message
----------------------------------------------------------------------------------------------------
[1250  ] [1   ] Loss: 1.8420 | PPL: 6.31 | GradNorm: 1.12 | 2950 tok/s | msg: Call me Ishmael.
[1251  ] [1   ] Loss: 1.8385 | PPL: 6.29 | GradNorm: 0.98 | 3012 tok/s | msg: Some years ago...

🤝 Contribution

Pull requests are welcome. For major changes, please open an issue first to discuss what you would like to change.

⚠️ Disclaimer

This tool is for educational and research purposes only.

  • Copyright: Please ensure you have the right to read/use the .epub files you load. The author does not endorse piracy.
  • Usage: The author is not responsible for any extended graduation dates, missed deadlines, or confusion caused by "exceptionally fast reading speeds" during working hours.

Happy "Training"! 🎓

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages