Skip to content
SIH Buddyby Ganeev Singh
Dev

πŸ”₯ Roast My Pick Β· SIH26165

AI/NLP Engine to Detect Serious Injury & Fatality (SIF) Precursors in OIL's Unsafe-Act/Unsafe-Condition and Near-Miss Reports

Oil India Limited

Mild33/100

Reasonable choice. The scoreboard liked it. The scoreboard is not the one asking questions on the day.

Worth considering. Well-specified and grounded in real safety science, but with no OIL data provided you train on constructed labels β€” build the Life-Saving Rule tagging against the published framework, handle the class imbalance deliberately, and be clear about what your classifier was validated on. Roughly 110–250 teams are expected to go here.

The receipts

Every red flag on this statement, in full. These are the four places it bites.

  1. Exhibit A

    OIL's actual reports are not provided and no public SIF-labelled safety corpus exists, so you construct or hand-label training data and the classifier's real validity depends on that

  2. It gets worse

    SIF potential is a subtle judgement even for human safety experts, so a classifier trained on limited labels will disagree with experts on the borderline cases that matter most

  3. Still reading?

    The SIF-potential class is a minority at roughly a fifth, so class imbalance must be handled or the classifier trivially predicts non-SIF

  4. And the finisher

    Misclassifying a genuine SIF-precursor as routine is exactly the failure the tool exists to prevent, so recall on the fatal-potential class matters more than overall accuracy

The damage report

Every score this statement earned, and what each one actually costs you.

  • Feasibility

    3/5

    Buildable. Not comfortably. There is a week in here you have not planned for yet.

    Text classification and tagging are standard NLP, and the SIF-precursor concept has published frameworks β€” IOGP Life-Saving Rules, the EEI model β€” to structure labels against, but OIL's actual reports are not provided and no public SIF-labelled safety corpus exists, so you must construct or hand-label training data, which bounds the classifier.

  • Innovation scope

    3/5

    Mildly interesting. The novelty will not carry the room; the build has to.

    The task is well-defined classification and tagging against established safety frameworks, so your room is in the SIF-potential modelling and precursor pattern extraction rather than in the concept, which the description specifies clearly.

  • Clarity

    5/5

    The ask is unambiguous, which quietly removes your favourite excuse.

    The description names the exact three outputs β€” SIF classification, Life-Saving Rule tagging, precursor pattern surfacing β€” cites the underlying safety-science frameworks, and even quantifies the roughly twenty to twenty-five percent SIF-potential proportion, making the requirement exceptionally clear.

  • Acceptance potential

    3/5

    Middle of the pack. This statement will not win the room for you β€” you will have to.

    The problem is well-specified, genuinely grounded in real safety science, and less crowded than generic NLP, but with no OIL reports provided you train on constructed or public-proxy safety text, so the classifier's real-world validity rests on data you assembled and an OIL safety judge will ask what it was trained and validated on.

  • Effort

    Heavy

    Heavy. Somebody on this team is not sleeping in week three. Pick who, on purpose.

    The classifier, the Life-Saving Rule tagger, the precursor pattern extraction and the density-ranking dashboard are focused, well-bounded pieces, with labelled-data construction the main effort.

  • Demo-ability

    Easy

    Easy to demo β€” and so is everyone else's. Working is the floor here, not the achievement.

    Separating the fatal-potential minority from the routine majority and ranking sites by risk density is a clear, immediately meaningful demo for any safety audience.

  • Data

    None supplied

    No dataset comes with this one, so every accuracy figure you quote is a number about labels you invented.

    Nothing is provided with the statement. You are sourcing, cleaning and labelling it yourself, and that work is invisible in the demo but very visible in the questions.

The demo they will have already seen

Somewhere around 110–250 teams are heading here, and the description is doing the choosing for most of them. They will read the same brief, reach the same architecture, and build a version of the same demo you are planning. Being correct is the floor. If your five minutes could be swapped with the team before you and nobody in the room would notice, you have not picked badly β€” you have built predictably, which costs exactly the same and hurts more.

What survives

The ground worth standing on when the questions start.

  • The SIF-precursor concept rests on published safety-science frameworks the description cites, so your labels and categories have an authoritative basis rather than being invented
  • Separating genuine fatal potential from routine reports is a real, high-value distinction that directly changes where a safety team acts
  • The Life-Saving Rules give you a fixed, well-defined tagging taxonomy to classify against

Nothing here is fatal. It is just the list of places this statement pushes back, and you now get to push there first.

The framing is a joke. The findings are not β€” they are the same analysis on the statement page, and every line above is attached to a score or a fact in the record. It is one opinion with its reasoning attached, so argue with it before you trust it.