About the role
Indian trade regulation moves constantly. A RoDTEP rate is halved in February and restored in March. A notification supersedes another notification that amended a third. Nobody sends you an email about it.
You build the system that watches, diffs, extracts and routes — so that a change in a CBIC circular becomes a machine-readable delta with an owner and a deadline, rather than something a customer discovers at the port.
This is the asset that compounds. Products can pivot and the wedge can move; a versioned, diffed, obligation-extracted corpus of Indian trade regulation survives all of it and is worth more every month.
One thing we will say plainly because it comes up: we are not looking for sentiment analysis. Regulation carries obligation, not opinion, and a sentiment score over a CBIC circular is both noise and unauditable. The techniques that matter here are change detection, obligation extraction, applicability mapping, event extraction and temporal reasoning. If you want to argue with that, do — it would be a good conversation.
This is not a general machine-learning role and it is not a model-training role. And it is not a role where the extraction is done when the demo works.
What you’ll do
- The versioned source register — DGFT notifications, CBIC circulars, tariff notifications, RoDTEP schedules, the Denied Entity List, and the long tail nobody indexes.
- Structural diffing. Not "this page changed" but "this clause changed, in this way, with effect from this date".
- Obligation extraction — turning prose into actor, modality, action, condition and deadline.
- Applicability mapping. Joining an obligation to a specific customer's HS codes, ports, destinations and scheme participation, so an alert is about them rather than about everyone.
- Event extraction over enforcement and trade news — typed events carrying entity, date, jurisdiction and citation.
- Citation networks and temporal queries. "What were the rules on the date of this shipment" is a different question from "what are the rules", and only the second one is easy.
What we look for
- You have built ingestion and extraction over documents that were never designed to be parsed, at a scale where you had to care about it.
- Change detection, entity resolution or information extraction somewhere in your history — in production, not in a notebook.
- You understand that provenance is the product. An extracted obligation without a citation and an effective date is worse than no obligation at all.
- You are comfortable with the unglamorous half: source discovery, scrapers that break, PDFs that are images, and the tenth different date format.
- Python, and enough infrastructure sense to run something that has to be right every day rather than once.
Nice to have
- Regulatory or legal text specifically — horizon scanning, compliance content, legal publishing, or a regtech product team.
- Knowledge graphs, temporal or bitemporal data modelling.
- Multilingual extraction.
- You have worked somewhere the extraction had to be defended to an auditor.