Skip to content
SIH Buddyby Ganeev Singh
Dev

πŸ”₯ Roast My Pick Β· SIH26056

Development of a Real-time Airfare Price Index for India through Automated Web Scraping of Airline and Online Travel Aggregator Portals for Augmentation of the Consumer Price Index (CPI).

MoSPI

Mild28/100

Reasonable choice. The scoreboard liked it. The scoreboard is not the one asking questions on the day.

Worth considering. The validation against a published official series is a rare and strong position, but the thirty-day back-test is a calendar requirement and the compliance tension is real, so start collecting immediately and make your missing-fare methodology the thing you defend rather than the scraper. Roughly 270–500 teams are expected to go here.

The receipts

Every red flag on this statement, in full. These are the four places it bites.

  1. Exhibit A

    The statement demands robots.txt and terms-of-service compliance while asking you to scrape sites whose terms typically prohibit exactly this β€” read the actual terms of your sources and be prepared to say which you excluded and why, because a judge from a statistics ministry will care about the legality more than the throughput

  2. It gets worse

    Thirty days of back-tested results is a calendar dependency, not an engineering one, so collection has to start on day one or the requirement is unmeetable regardless of how good the code is

  3. Still reading?

    A sold-out or unavailable flight is not a high price and treating it as one silently biases the index upward β€” how you handle missingness is the single most consequential methodological choice here

  4. And the finisher

    Anti-bot defences change without warning and a collection engine that worked last week can silently return empty results, so build monitoring that fails loudly rather than a scraper that quietly stops

The damage report

Every score this statement earned, and what each one actually costs you.

  • Feasibility

    3/5

    Buildable. Not comfortably. There is a week in here you have not planned for yet.

    The validation target is genuinely public since monthly average fare data is published, and the cleaning and index construction are ordinary work, but the collection layer sits in real tension with itself β€” the statement demands robots.txt and terms-of-service compliance from sites whose terms generally prohibit automated fare collection, and those sites deploy active anti-bot defences precisely to enforce that.

  • Innovation scope

    3/5

    Mildly interesting. The novelty will not carry the room; the build has to.

    The pipeline shape, the city pairs, the advance-purchase windows and the deliverables are all specified, but the index methodology itself β€” how you weight routes, handle missing quotes and treat a fare that is unavailable rather than expensive β€” is genuine statistical design and is where the substance of this submission lives.

  • Clarity

    5/5

    The ask is unambiguous, which quietly removes your favourite excuse.

    Exceptionally specific for a statistics statement β€” it names the airlines, the aggregators, example city pairs, all five advance-purchase windows, the fare components to separate, the anti-bot problems to handle and a concrete thirty-day back-test requirement against a named public benchmark.

  • Acceptance potential

    3/5

    Middle of the pack. This statement will not win the room for you β€” you will have to.

    Real public validation data, a sharply specified brief and an official statistics domain that almost no team will enter all help, but the collection layer's compliance tension is structural rather than solvable, and thirty days of back-tested results is a calendar requirement that a team starting late simply cannot satisfy.

  • Effort

    Heavy

    Heavy. Somebody on this team is not sleeping in week three. Pick who, on purpose.

    A resilient multi-source collection engine, a cleaning pipeline, an index construction module, a dashboard, an API and thirty days of actual back-tested collection is five components plus a calendar-time dependency, and scrapers against defended sites need continuous maintenance rather than one-time construction.

  • Demo-ability

    Medium

    Demoable, if you rehearse it. Nobody rehearses it.

    The back-test overlay is genuinely persuasive because it validates against an external published series, and the lead-time elasticity curve is a striking visual, but the product is charts and an API rather than anything interactive.

The demo they will have already seen

Somewhere around 270–500 teams are heading here, and the description is doing the choosing for most of them. They will read the same brief, reach the same architecture, and build a version of the same demo you are planning. Being correct is the floor. If your five minutes could be swapped with the team before you and nobody in the room would notice, you have not picked badly β€” you have built predictably, which costs exactly the same and hurts more.

What survives

The ground worth standing on when the questions start.

  • Published monthly average fare data gives you an external validation target, so you can show your index tracks an independent official series rather than merely asserting it is correct β€” very few statements offer that
  • The index methodology is the real intellectual content and it is genuinely open: how you weight routes and how you treat an unavailable fare are decisions a statistician will interrogate and reward
  • Official statistics is a domain essentially no student team enters, so the field for this will be very thin

Nothing here is fatal. It is just the list of places this statement pushes back, and you now get to push there first.

The framing is a joke. The findings are not β€” they are the same analysis on the statement page, and every line above is attached to a score or a fact in the record. It is one opinion with its reasoning attached, so argue with it before you trust it.