What if a machine's sensor readings could actually change its owner's loan terms — not just predict a breakdown, but say something about whether the loan gets repaid?
Built for the Razorpay AI Buildathon (Open Track), by a mechanical engineering student.
Watch here: https://www.youtube.com/watch?v=jejCNaZvNN8
Razorpay's Open Track basically says: build something that doesn't fit a category, as long as it solves a real problem and actually works. That's this project. It lives right between two things — mechanical engineering (machine health, failure modes, downtime) and fintech (lending, risk, repayment) — and I picked it because I actually understand the mechanical side, not because it sounded impressive on paper.
Here's something that bugged me once I looked into it: when a bank or NBFC lends money against a machine — a lathe, a CNC setup, whatever — they price that loan once, at the start, based on credit score and the machine's age. After that, nobody checks on the machine again. If it starts failing, the lender has no idea until a payment gets missed.
A few numbers that convinced me this is real, not just a nice-sounding idea:
- Unplanned downtime costs Indian MSME manufacturers anywhere from ₹50,000 to ₹5,00,000 an hour
- Not having spare parts on hand causes up to half of all MSME downtime
- Indian SMEs take about 73 days on average just to collect payment on invoices, and manufacturing credit cycles often stretch 60–120 days
- Machinery loans (NBFC/bank) typically run ₹10L–1Cr at 12–18% interest, priced once and never revisited
- I looked — there's no published research anywhere connecting real-time machine sensor data to whether a loan gets repaid. That link just doesn't exist yet.
I'm not claiming to have solved that last point. What I built is a first, honest step toward it — turning machine health into a cash-flow risk signal, and being upfront about exactly which parts of that are proven and which parts are still assumptions.
I didn't want to just train one model and call it "AI predicts loan default," because that's not something I can back up with data. So instead I broke it into four separate steps, and only used AI where it actually made sense:
- Sensor readings → failure prediction. This is the one real ML model in the whole thing — trained on actual data, with real accuracy numbers I can show you.
- Failure type → how many hours of downtime. Just arithmetic, based on a downtime estimate I looked up and can justify (see
ASSUMPTIONS.md), and it's adjustable in the demo, not baked in. - Downtime → rupees at risk. Also arithmetic, using the downtime-cost figures above.
- Rupees at risk vs. the loan's monthly EMI → a risk flag. A simple threshold rule, also adjustable.
On top of that:
- One LLM call turns those four numbers into a plain-English note, like something a loan officer could actually read. That's the only place an LLM touches this pipeline — everything else is a trained model or plain math.
- Two Razorpay actions — a Payment Link if something looks risky, a Payout if a machine looks healthy enough to deserve more credit. Both require a manual approval click before they actually fire, and every decision gets logged.
I compared three different modeling approaches instead of just picking one and hoping:
| Model | Macro F1 |
|---|---|
| Random Forest (class-weighted) | 0.53 |
| XGBoost | 0.57 |
| Random Forest + SMOTE (the one I kept) | 0.585 |
I used macro F1 instead of plain accuracy because accuracy alone barely notices rare failure types — macro F1 treats every failure type as equally important, which felt like the honest way to judge this.
The winning model did well on Heat Dissipation Failures (F1 0.88) and got noticeably better at Tool Wear Failures once I added oversampling (0.10 → 0.24).
One thing I want to flag rather than bury: Random Failure scored 0.00 across every single model I tried. At first that looked like a bug. It isn't — "Random Failure" in this dataset is literally defined as failure with no sensor pattern behind it. No model should be able to predict something that's random by definition, so getting 0.00 there is actually the model working correctly, not failing.
I also ran the whole pipeline on 100 real cases at once (not hand-picked ones) — 98 came out healthy, 2 came out flagged as cash-flow risk, and the model was 100% accurate on that batch.
Then I checked something I cared about: does my "only flag it if it actually threatens the loan payment" logic do anything different from just flagging every predicted failure? Honestly — on this batch, no, both approaches flagged the same 9 out of 300 cases. I dug into why instead of hiding it: it turns out my model's strongest real predictions are Heat Dissipation Failures, and the downtime cost for those (around ₹6.75 lakh) is high enough to matter regardless of how big the loan is, at least within the loan sizes I tested. That's not a failure of the idea — it's a real finding: some failures are serious enough that no amount of "but it's a big loan" changes the picture.
I'm listing these because a smooth story usually means someone's leaving stuff out, and I'd rather be honest:
- XGBoost refused to run because my sensor columns had names like "Air temperature [K]" — turns out XGBoost doesn't like square brackets in column names. Fixed by renaming columns just for that one model.
- Ran out of disk space mid-install, even though my project was on a different drive with plenty of room — pip was quietly using my C drive for its cache and temp files. Redirected those to the same drive as the project.
- Went through three different AI providers before one worked. Claude's API account had zero credits. Google's Gemini kept throwing an authentication error that didn't go away even after switching to their newest SDK. Groq finally worked — after I fixed one more thing, which is that Groq had quietly deprecated the model I first tried, so I had to swap in their current one.
- Couldn't finish Razorpay's business onboarding because it asks for a PAN, and I don't have one. I wasn't going to fake KYC details on a real financial platform, so instead I built the integration to run against the real Razorpay SDK, and it falls back to realistic simulated responses when there's no live account behind it. The code is real and correct; only the actual network call is standing in for now.
I want to be clear about the difference between what I built and what would actually validate it:
- The real test would be a pilot with an actual lender's repayment data, checking whether this risk signal correlates with what really happens over six months or a year. That's the one thing that would turn this from a hypothesis into something provable.
- The downtime-hour estimates I used are reasonable engineering judgments, not measured repair-time data — swapping in real OEM or service-log numbers would make Link 2 much stronger.
- The Razorpay side just needs KYC completed — the code doesn't need to change.
- Right now this only works for one sensor dataset and one kind of machine. A real version would need models per machine type.
pip install -r requirements.txt
python generate_synthetic_data.py # only a placeholder until you have the real dataset
# Get the real one: search "AI4I 2020 Predictive Maintenance Dataset" on Kaggle,
# save it as data/ai4i2020.csv (see README_DOWNLOAD_DATA.md)
python model_comparison.py # trains all three models, picks the best, prints real numbers
python batch_test.py
python naive_vs_smart_flagging.py
# Pick one LLM provider (it checks in this order: Anthropic, then Groq, then Gemini)
$env:ANTHROPIC_API_KEY="sk-..." # or
$env:GROQ_API_KEY="gsk_..." # or
$env:GEMINI_API_KEY="AIza..."
# Razorpay is optional -- runs honestly in mock mode without it
$env:RAZORPAY_KEY_ID="rzp_test_..."
$env:RAZORPAY_KEY_SECRET="..."
streamlit run app.py
machinescore/
├── README.md <- this file
├── ASSUMPTIONS.md <- where every number came from
├── PITCH_SCRIPT.md <- what I say in the pitch video
├── README_DOWNLOAD_DATA.md <- how to get the real dataset
├── requirements.txt
├── generate_synthetic_data.py <- placeholder data generator, not the real thing
├── train_model.py <- a simpler, single-model version
├── model_comparison.py <- compares three models, real metrics
├── risk_chain.py <- downtime -> money -> flag logic
├── batch_test.py <- runs the whole thing on 100+ cases at once
├── naive_vs_smart_flagging.py <- checks whether my flagging logic actually helps
├── llm_explain.py <- the one LLM call, three providers supported
├── razorpay_actions.py <- the two Razorpay actions, approval gate, mock mode
├── app.py <- the one-screen demo
├── data/ <- ai4i2020.csv goes here
└── model/ <- trained model and results land here