Skip to content
SIH Buddyby Ganeev Singh
Dev
All problem statements
SIH26103Strong pickacceptance 4/5

Use case on web-based integrated project-monitoring platform

MoSPI · Smart Automation · Software

Real public data, a directly computable target and a sponsor thoughtful enough to ask whether AI is even the right tool — answer that question honestly and tell them which missing fields would improve prediction, because that is a finding they can act on and almost nobody else will produce it.

Open dataset ↗

Data: The Project Monitoring Report for April 2026 may be referred to for the key fields and parameters: https://paimana-proj.mospi.gov.in/ReportPage

What it actually is

Nearly two thousand large central infrastructure projects are tracked monthly, and a great many of them end up costing far more and taking far longer than approved. The monitoring system records all of it faithfully but only after the fact. The ask is to use twenty years of that record to predict which projects are heading for overruns before they get there.

What to build

A predictive monitoring layer over the project database, with the target variables directly computable from published fields since cost overrun is revised cost against original and time overrun is actual against scheduled completion: models predicting each at a project's current stage from the attributes available at that stage, a project-level risk score ranking the portfolio by likely trouble, driver analysis identifying which factors actually move the outcome, and an early warning view for administrators — plus the two research questions the statement explicitly poses and that most teams will skip, namely whether machine learning genuinely outperforms conventional statistical methods here, and how much of the predictive power comes from the fields currently captured versus variables the monitoring form does not collect at all, which is a finding the ministry could act on directly.

Smallest thing that wins the room

Take projects as they stood at an earlier date, predict which would overrun, and show the ranked list against what actually happened since — with the comparison of your model's accuracy against a straightforward regression baseline beside it.

How crowded this one gets

A guess, projected from the 2025 statements — the last year where both the submission counts and the winners were published.

Moderate120–270 teams expectedroughly 1 in 99–230 wins it

Quieter than 54% of the 226 · #105 of 226 by expected field

A normal-sized field. Your idea has to be good, not miraculous.

Why: central ministry statements sat below the average.

This is a guess, not a fact

Nobody has published 2026’s numbers yet. This is an analysed estimate from last year’s pattern, so please do not take it as the truth — check the live counter on the SIH portal before you decide anything. The range covers the middle half of likely outcomes, so one statement in two lands outside it. Entry closes at 500 ideas per statement, so no range goes past that — a statement that reaches the cap fills and shuts rather than drawing an unlimited crowd. The model reads only three things a team can see before choosing — software or hardware, the theme, and what kind of body posted it — and those explain about a quarter of the variation in last year’s field sizes (R² 0.25 on held-out statements). Trust the band more than the number, and the ordering more than either. It cannot see how good your idea is, which is the part that actually decides it.

The scores

The number is the shorthand. The line under it is the reason.

What you will be writing

  • survival analysis for time-to-completion overrun
  • gradient boosting cost overrun regression
  • stage-aware feature construction from monitoring panel
  • conventional statistical baselines for comparison
  • SHAP driver attribution for cost escalation
  • public PAIMANA monitoring report ingestion
  • Infrastructure project monitoring
  • Predictive analytics for governance
  • Public expenditure management

Prior art to read before you start

cost and schedule overrun prediction · project risk scoring and early warning · predictive versus conventional method comparison

Analysed by Claude Opus. Every score above is a judgment call with its reasoning attached — kindly cross-check this against the official statement on the SIH portal before your team commits to it.