Skip to content
SIH Buddyby Ganeev Singh
Dev

πŸ”₯ Roast My Pick Β· SIH26122

Intelligent Data Capture & Schedule-Linking Layer for Infrastructure Project Management: Real-Time Actual Progress Tracking (Planning-to-Execution Bridge)

Oil India Limited

Mild17/100

Good pick. Genuinely. Now sit down, because the judges are going to try anyway β€” and this is what they will try.

Strong pick. The semantic matching is the whole submission and it is genuinely interesting work in a field nobody else will enter β€” build the confidence threshold and reviewer queue from the start, because a silent wrong match is the failure mode that matters. Roughly 85–190 teams are expected to go here.

The receipts

Every red flag on this statement, in full. These are the four places it bites.

  1. Exhibit A

    No paired real-world schedule and site report dataset exists publicly, so you construct both sides yourself and your matching accuracy is measured on data you authored

  2. It gets worse

    Field execution is often finer-grained than the plan, so many reports map to a fraction of an activity rather than to one cleanly β€” handling partial progress is the hard case and teams tend to skip it

  3. Still reading?

    A wrong automatic match silently corrupts the schedule, which is worse than no match, so the confidence threshold and reviewer queue are core requirements rather than polish

  4. And the finisher

    The description also asks for a closed-project knowledge base feeding future planning, which is a second product that teams will quietly drop

The damage report

Every score this statement earned, and what each one actually costs you.

  • Feasibility

    3/5

    Buildable. Not comfortably. There is a week in here you have not planned for yet.

    The semantic matching core is achievable with sentence embeddings over activity descriptions, and Primavera XER and MS Project XML are documented formats you can parse, but no real project's paired reports and schedules are public so you must construct a realistic work breakdown and report set yourself.

  • Innovation scope

    4/5

    There is something genuinely new here. Do not bury it under another dashboard.

    Matching loosely worded field language to structured activity IDs across disciplines that describe the same work differently is a genuinely open problem, and how you handle field granularity being finer than the planned breakdown is entirely your design.

  • Clarity

    4/5

    The ask is unambiguous, which quietly removes your favourite excuse.

    The description states the exact failure precisely, gives the worked example of 'spool erected' versus 'Erect Line 24-XX', and enumerates the required inputs and outputs, though it does not say how accurate the matching must be.

  • Acceptance potential

    4/5

    Strong footing before you have written a line. Try not to waste it.

    A quiet gem β€” project controls is unglamorous enough that almost nobody will pick it, the semantic matching core is real technical work rather than another dashboard, and the pain is instantly recognisable to any judge who has run a capital project.

  • Effort

    Heavy

    Heavy. Somebody on this team is not sleeping in week three. Pick who, on purpose.

    Multi-format ingestion, the matching engine, the reviewer interface and the schedule write-back are four pieces, and constructing a believable project schedule with matching reports is itself substantial preparation.

  • Demo-ability

    Medium

    Demoable, if you rehearse it. Nobody rehearses it.

    The match-and-update moment is clear and satisfying, but a judge needs the planning context explained before they understand why linking a sentence to an activity ID matters.

  • Data

    None supplied

    No dataset comes with this one, so every accuracy figure you quote is a number about labels you invented.

    Nothing is provided with the statement. You are sourcing, cleaning and labelling it yourself, and that work is invisible in the demo but very visible in the questions.

The demo they will have already seen

Somewhere around 85–190 teams are heading here, and the description is doing the choosing for most of them. They will read the same brief, reach the same architecture, and build a version of the same demo you are planning. Being correct is the floor. If your five minutes could be swapped with the team before you and nobody in the room would notice, you have not picked badly β€” you have built predictably, which costs exactly the same and hurts more.

What survives

The ground worth standing on when the questions start.

  • The description hands you the exact matching example, so you know precisely what the hard case looks like and can build and demonstrate against it
  • The semantic matching core is defensible technical work that separates this from the many project dashboards a judge will see
  • A voice reporting path for supervisors is explicitly invited and is a small addition that addresses the real reason field data is poor

Nothing here is fatal. It is just the list of places this statement pushes back, and you now get to push there first.

The framing is a joke. The findings are not β€” they are the same analysis on the statement page, and every line above is attached to a score or a fact in the record. It is one opinion with its reasoning attached, so argue with it before you trust it.