Skip to content
SIH Buddyby Ganeev Singh
Dev

πŸ”₯ Roast My Pick Β· SIH26167

SatQuery AI - An Interactive Vision-Language Assistant for Multimodal Remote Sensing Image Analysis through Text Queries

Indian Space Research Organisation(ISRO)

Mild34/100

Reasonable choice. The scoreboard liked it. The scoreboard is not the one asking questions on the day.

Worth considering. The datasets are handed to you and the optical-SAR angle is genuinely valuable, but VLM fine-tuning is compute-heavy and the multisensor reasoning is what tends to get dropped β€” commit to the paired and cross-sensor cases, since a single-image VQA model is the crowded easy version. Roughly 90–210 teams are expected to go here.

The receipts

Every red flag on this statement, in full. These are the four places it bites.

  1. Exhibit A

    VLM fine-tuning is genuinely compute-heavy, and without adequate GPU access you will ship something under-trained

  2. It gets worse

    The multisensor and multitemporal reasoning is the hard, distinguishing part and is exactly what teams drop under time pressure, leaving a single-image VQA model

  3. Still reading?

    A domain-adapted VLM can still hallucinate confident wrong answers about imagery, which for an ISRO judge is worse than an honest refusal

  4. And the finisher

    The scope spans adaptation, fusion, change reasoning and multi-benchmark evaluation, so shallow breadth is a real risk

The damage report

Every score this statement earned, and what each one actually costs you.

  • Feasibility

    3/5

    Buildable. Not comfortably. There is a week in here you have not planned for yet.

    BigEarthNet, VRSBench and RSVQA are named public datasets and remote-sensing VLM fine-tuning is an active area with base models to build on, but fine-tuning a VLM to reliably handle multisensor and multitemporal reasoning is compute-heavy and genuinely hard, and doing it well within a hackathon is ambitious.

  • Innovation scope

    4/5

    There is something genuinely new here. Do not bury it under another dashboard.

    Multisensor optical-SAR fusion and multitemporal reasoning through a single language interface is genuinely open β€” most remote-sensing VLM work handles single optical images, so the paired and cross-sensor reasoning is where the real contribution lies.

  • Clarity

    4/5

    The ask is unambiguous, which quietly removes your favourite excuse.

    The description names the datasets, the required capabilities including multisensor and multitemporal reasoning, the adaptation requirement and the evaluation benchmarks, so the deliverable is well defined despite being broad.

  • Acceptance potential

    3/5

    Middle of the pack. This statement will not win the room for you β€” you will have to.

    The datasets and benchmarks are named and the multisensor angle is genuinely valuable, but VLM fine-tuning is compute-heavy and the scope is large, so a team that ships a single-image RS-VQA model has answered the easy part while the multisensor and multitemporal reasoning that distinguishes this is exactly what tends to get dropped.

  • Effort

    Massive

    A semester of work wearing a hackathon costume. Something is getting cut; decide what now, not in week five.

    Domain-adapting a VLM, handling multisensor and multitemporal inputs, possibly orchestrating specialist models, and evaluating on multiple benchmarks is a large, compute-intensive undertaking.

  • Demo-ability

    Easy

    Easy to demo β€” and so is everyone else's. Working is the floor here, not the achievement.

    Asking a plain-language question and getting a grounded answer about a satellite image is immediately compelling, and the change-detection and optical-SAR cases make strong demo beats.

The demo they will have already seen

Somewhere around 90–210 teams are heading here, and the description is doing the choosing for most of them. They will read the same brief, reach the same architecture, and build a version of the same demo you are planning. Being correct is the floor. If your five minutes could be swapped with the team before you and nobody in the room would notice, you have not picked badly β€” you have built predictably, which costs exactly the same and hurts more.

What survives

The ground worth standing on when the questions start.

  • BigEarthNet, VRSBench and RSVQA are named public datasets, so both training and evaluation data are handed to you
  • The optical-SAR fusion and multitemporal angles are genuinely differentiating in a field where most RS-VLM work is single-image optical
  • Natural-language querying of satellite imagery is an immediately compelling and legible demo

Nothing here is fatal. It is just the list of places this statement pushes back, and you now get to push there first.

The framing is a joke. The findings are not β€” they are the same analysis on the statement page, and every line above is attached to a score or a fact in the record. It is one opinion with its reasoning attached, so argue with it before you trust it.