Hashim

Product Owner + Engineer

WorkAboutResume

Daftarkhwan  ·  Product Owner  ·  2026

Building An Operations System For Fifteen Buildings And 10,000+ Assets.

Turning fifteen WhatsApp threads into a structured operations platform — incident reporting, asset-linked repair history, derived priority and routing, blocker attribution, preventive maintenance and repair-versus-replace analysis.

Before · fifteen threadsBranch 07“AC is making noise again”12Branch 02“generator tripped”47Ops · Escalations“members complaining”61Branch 07“AC stopped cooling”8What none of it left behind×asset history×a known owner×a derived priority×a blocker×a cost trailAfter · one asset recordAC-034 · Room ABranch07CategoryHVACPriorityP2 — derivedOwnerHVAC teamSLAclock runningPrevious repairs4Current blockerFinancePKR 110,000 spent on this one unit in two monthsEvery line above was already in the messages
The messages already contained the facts. The product made them accumulate against the unit instead of scrolling away.
15 → 1

Branch threads replaced by one queue

5 → 2

Steps between a fault and work starting

10,000+

Assets with a real repair history

The Engagement

Turning fragmented communication into operational memory.

Daftarkhwan operates roughly fifteen coworking locations across Pakistan, serving thousands of members and managing more than ten thousand physical assets — HVAC, electrical, furniture, internet, plumbing and general facilities.

Operational coordination lived primarily inside WhatsApp. Branch conversations, operations threads and repair discussions created activity without creating durable operational memory.

I worked with the team to turn that into a structured operations platform: incident reporting, asset-linked repair history, automatic priority and routing, blocker attribution, preventive maintenance, repair-versus-replace analysis, and selective use of AI.

The goal was not to build a ticketing system. It was to give the organisation operational memory.

Client Constraint

Every repair created more messages, but almost no organisational memory.

Facilities teams were repairing air conditioners, generators, internet connections and furniture every day. The problem was how that work was recorded.

A fault could begin in one WhatsApp group, move into operations, get escalated to facilities, be discussed with finance, and disappear into chat history once fixed. The next time the same asset failed, it looked like a new problem.

Branch 04Ops · escalationsBranch 11Branch 07FacilitiesBranch 02Branch 1310,000+ assets being remembered by peopleOne asset systemHistoryRoutingOwnershipBlockersSLA
Fourteen branch threads plus an operations thread. Every fault, escalation and follow-up lived in one of them, and the record of what had already been done lived nowhere.

That made prioritisation, routing, recurring-fault detection, maintenance planning, repair-cost accumulation and performance measurement unnecessarily difficult. The business did not lack information. It lacked structured history.

In chat · four unrelated messagesJanuary“AC is leaking”March“AC making noise again”June“AC stopped cooling”August“AC issue in Room A”Four incidents. No link between any of them.Against asset AC-034 · one recurring failureJanuaryrepair 1Marchrepair 2Junerepair 3Augustrepair 44 repairs · 2 months · PKR 110,000 · one unit
The same unit, four times. On the left it is four conversations; on the right it is one failing compressor with a cost attached.

The organisation remembered incidents as conversations instead of remembering them as events attached to assets.

System Hypothesis

Turn every facilities issue into a structured record attached to the asset.

The first solution was to replace fragmented reporting with one operational queue. A member or employee reports an issue, and the system creates a structured ticket — title, description, priority, assignee, due date, asset, location, status, owner and activity history.

Each incident enters the system, attaches to a physical asset, moves through a defined workflow, retains its history and contributes to future maintenance decisions.

That created the data model required for

  • Asset history
  • Recurring-fault detection
  • Maintenance scheduling
  • Repair spending
  • Performance analysis
  • Replacement recommendations

If every repair were captured properly, the company could finally reason across time instead of managing one incident at a time.

What Broke In Deployment

The software worked. The workflow assumptions didn’t.

The first version behaved like a conventional ticketing system. It asked reporters for priority, assignee and due date. All three produced poor operational data.

Self-reported priorityP1P2P3P4So priority stopped belonging to the reporter
Shape, not measured counts — this is the distribution operations described. Marking your own problem urgent is the rational move when urgency is self-declared.

Priority

Almost everyone marked their own problem urgent.

Assignee

Reporters often did not know which technician, department or vendor owned a fault.

Due date

Asked when something should be fixed, users naturally chose today.

Performance

A technician could finish diagnosis quickly, then wait days for finance, procurement, a vendor, parts or access. The ticket showed total elapsed time, not who owned the delay.

The common thread was not poor user behaviour. The system was asking people questions they were not qualified or incentivised to answer accurately.

Before · five stepsno repair work happens hereFaultreportedOps readsOps forwardsTech opensWorkstartsTwo people were moving information, not solving the faultAfter · two stepsMember identifies the assetWork starts
The two shaded points are hand-offs. Nobody standing on one of them was working on the fault; they were reading it in order to give it to someone else.

The first version digitised the existing process. It had not yet redesigned it.

Redesigning The Operating Model

I removed decisions from people when the system could derive them from the asset.

The breakthrough was separating what the reporter genuinely knew from what the organisation already had. The reporter really knew only two things: what is broken, and what the problem looks like. Most of the remaining workflow could be derived.

Priority came from the asset

I defined rules based on asset and failure type, so urgency stopped being a claim and became a property of the equipment.

Routing came from the asset

Branch, category, department, technician pool and ownership determine where the ticket goes.

The form shrank

The user primarily needs to identify the asset and describe the fault. Priority and assignee came off the form entirely.

The human supplies · two things they cannot get wrongAC-034 · Room A“Making noise, not cooling”The asset supplies · already known, never asked forBranchCategoryDepartmentPriority classTechnician poolThe system derives · nobody types any of thisRoutingOwnerPrioritySLAEscalationPriority and assignee were deleted from the form entirelyRules are editable by the client — changing them does not need us
Two inputs, and everything below them already known or derivable. The reporter is never asked a question they have an incentive to answer badly.

Delays attach to blockers

The workflow tracks technician diagnosis, finance approval, vendor response, procurement and repair work separately, so elapsed time lands on whoever owns it.

What the report used to showEngineer A · 21 days open · 30% completionManagement conclusion: underperformingWith the blocker recorded on the same ticketInspectedBlocked · Finance · funds not releasedFixedday 0day 21Engineering-owned time · 8 daysFinance-owned time · 13 daysManagement conclusion: the work stalled outside engineering
The same twenty-one days, split by who actually owned them. An illustrative single ticket — the point is the split, not the days.

Maintenance became asset-centric

Each repair accumulates against the asset, exposing repair frequency, recurring faults, cumulative cost, maintenance cadence and time between failures.

AI only where interpretation helped

Rules handle priority, routing, SLA, escalation and permissions. AI handles title generation, repair-history summarisation, recurring-fault interpretation and explaining repair-versus-replace recommendations.

Rules · deterministicPriorityRoutingSLAEscalationPermissionsAI · interpretationDescription → titleFault interpretationNote summarisingPattern detectionHistory summariesHuman · judgementSpend approvalReplacement callsAmbiguous casesPolicy exceptionsDeterministic where the answer must be predictableAI only where interpretation creates leverage — and never on spend
The dividing line held from the first version onward. Rules where the answer must be predictable, a model only where interpretation earns something, and the money decision left with a person.

Rules where the answer must be predictable. AI where interpretation adds value. Humans where judgement still matters.

Operating Impact

From fifteen WhatsApp threads to one operational record across 10,000+ assets.

The system changed both the workflow and the decisions the organisation could make.

15 → 1

Branch threads replaced by one operational system

5 → 2

Steps between a fault and work starting

10,000+

Assets with persistent repair history

  • Meaningful priority, derived from operational context rather than self-declared.
  • Performance measured against blockers, not just assignees.
  • Preventive maintenance enabled by asset history.
  • Repair-versus-replace decisions supported by repair frequency, cumulative spend, projected cost and replacement cost.
BeforeAfter
FaultsDisappear into chatBecome asset history
PriorityChosen by the reporterDerived from the asset
RoutingOps interprets every requestDerived from the asset
Steps to workFiveTwo
DelayBlamed on the assigneeAttributed to the blocker
MaintenanceReacts after failureScheduled against the record
Repair vs replaceA judgement callCosted against history
Routing rulesWould need a developerOwned by the client
Asset AC-034 · 4 repairs · 2 months · PKR 110,000 already spentKeep repairing · projected 12 monthsPKR 400,000warranty · nonerepeat-failure risk · highReplace oncePKR 200,000warranty · 1 yearrepeat-failure risk · lowerCalculated by the deterministic layer from repair historyAI explains the recommendation. A human approves the spend.Structured history created this evidence, not a model
One asset, drawn to scale. The arithmetic is deterministic; the model explains it; a human signs it off.

AI did not create the evidence. Structured asset history did.

AI could explain the recommendation. The underlying decision came from reliable operational data.

What This Taught Me About Operational Software

The biggest improvement came from asking the system to know more, and asking the user to decide less.

Do not ask users for data they benefit from distorting.

Self-declared priority was a product-design problem, not just a data-quality problem.

Reduce forms by moving intelligence into the data model.

Better asset context allowed priority and routing to be derived automatically.

Measure the blocker, not just the person.

A single completion metric can hide where time is really being lost.

Structured history creates leverage later.

Repair-versus-replace analysis, preventive maintenance and AI interpretation were only possible because every incident was attached to an asset.

Use AI after the system has reliable context.

Models are powerful at interpretation and poor substitutes for missing structure.

Where each kind of decision lives

  • Deterministic logic — priority, routing, SLA, escalation, permissions.
  • AI — interpretation, summarisation, pattern explanation, ambiguous cases.
  • Humans — spending decisions, replacement approval, policy exceptions, high-judgement cases.

The real product was not a better ticketing system. It was an operational memory layer that let the company make better decisions about its buildings, assets and teams.

FAQ

The questions this project usually gets.

Why didn’t you just build a normal ticketing system?

That was my first approach. It worked technically, but the assumptions were wrong. I asked users to choose priority, assignee and due date, and those fields quickly became unreliable — everyone thought their own issue was urgent, most reporters didn’t know who owned the fault, and due dates became “today”.

The redesign was less about better software and more about deciding what the user should provide versus what the system could derive itself.

Why did you make the asset the centre of the system?

Because tickets disappear once they’re closed, but the asset stays. Once every repair was attached to the same physical unit, I could build a real history: how often it failed, what had already been spent on it, how long repairs lasted, and whether the same issue kept recurring.

That one architecture decision unlocked preventive maintenance and repair-versus-replace later.

Where did AI actually fit into this?

I deliberately kept AI out of anything that had to be predictable. Priority, routing, SLA, escalation and permissions were deterministic.

AI was useful for interpretation — cleaning up issue descriptions, summarising repair history, surfacing recurring patterns and explaining repair-versus-replace recommendations.

Rules when the answer must be consistent, AI when interpretation helps, and humans when judgement or money is involved.

What happened when the real-world process got messy?

That was most of the project. Information started in WhatsApp, ownership crossed facilities, operations, finance, vendors and procurement, and tickets could sit open for reasons that had nothing to do with the technician.

I had to model those blockers explicitly instead of pretending every ticket moved cleanly from open to done. That also stopped the system judging people by a misleading total-completion-time metric.

How did you make sure the system could still work if AI failed?

The core workflow did not depend on AI. A fault could still be reported, routed, prioritised, escalated and closed using the structured data and rules. AI sat on top as an assistive layer.

If an AI feature failed or gave a weak answer, the operational system still functioned. That separation was important for reliability.

How did you explain all of this to non-technical teams?

I avoided architecture language. Instead of “assignment is derived from asset metadata”, I’d say:

The reporter should only tell us what they know — what’s broken and what it looks like. The system should already know where the asset is and who owns it.

That made it much easier for operations and facilities to challenge the workflow and help shape the rules.

What was the biggest trade-off you made?

Not automating everything. There were plenty of places where I could have added AI, but doing so would have made the system harder to trust.

I kept high-cost decisions like replacement approval and policy exceptions with humans, while automating repetitive decisions the data could support reliably.

The goal was not maximum automation; it was the right division of responsibility between software, AI and people.

What would you do differently if you rebuilt it today?

I’d add observability earlier — dashboards for routing failures, manual overrides, bad asset metadata, blocker frequency, AI confidence and recommendation acceptance.

I’d also put more explicit evals around the AI layer, so I could measure when summaries or recommendations were useful versus when users corrected them.

Production AI is as much about monitoring and fallbacks as it is about model quality.

An eight-panel illustrated summary: fifteen WhatsApp threads with no accumulating history; the same air conditioner fault recurring under asset AC-034; the proposed single queue with asset-linked incidents; what worked; the three fields that failed in version one; what changed after the audit; where rules, AI and humans each decide; and what shipped.