-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathatom.xml
More file actions
909 lines (868 loc) · 108 KB
/
Copy pathatom.xml
File metadata and controls
909 lines (868 loc) · 108 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
878
879
880
881
882
883
884
885
886
887
888
889
890
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom"><title>Mohit Ranka</title><link href="https://www.mohitranka.com/" rel="alternate"/><link href="https://www.mohitranka.com/atom.xml" rel="self"/><id>https://www.mohitranka.com/</id><updated>2026-07-21T10:00:00+05:30</updated><subtitle>Senior engineering leadership for hard, well-defined problems.</subtitle><entry><title>The line manager role is disappearing faster than you think</title><link href="https://www.mohitranka.com/blog/line-manager-role-is-disappearing/" rel="alternate"/><published>2026-07-21T10:00:00+05:30</published><updated>2026-07-21T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2026-07-21:/blog/line-manager-role-is-disappearing/</id><summary type="html"><p>The line manager role is disappearing faster than most career plans assume. Meta has already moved toward one senior manager for roughly every fifty developers. That is not a distant forecast. It is an operating choice already in the market. If you lead engineers — or want to — the useful question …</p></summary><content type="html"><p>The line manager role is disappearing faster than most career plans assume. Meta has already moved toward one senior manager for roughly every fifty developers. That is not a distant forecast. It is an operating choice already in the market. If you lead engineers — or want to — the useful question is not “will this happen?” It is <strong>what kind of leader still has a job when the org chart gets that thin.</strong></p>
<figure class="post-figure">
<img src="https://www.mohitranka.com/images/blog/line-manager-role-is-disappearing.jpg" alt="Abstract illustration of a flattened org lattice with AI agent nodes around a central leadership figure" loading="lazy" width="1200" height="675">
<figcaption>
<span class="fig-caption">Flatter orgs do not remove leadership — they remove room for leaders who only schedule work.</span>
</figcaption>
</figure>
<!--more-->
<h2>Where this is going</h2>
<p>I see three moves landing at once:</p>
<p><strong>A. AI agents write most of the code.</strong> One senior developer directing a swarm of agents is not science fiction. It is the near-term shape of delivery: humans own intent, constraints, and review; machines own volume. Span of control for <em>output</em> rises even when headcount does not.</p>
<p><strong>B. The middle manager who mainly runs one-on-ones and tracks tickets will struggle to justify the seat.</strong> Coordination theater does not survive a world where execution bandwidth is cheap and coordination tools are better. If your value is status meetings and a dashboard of green/yellow/red, you are competing with software that does that without a skip-level.</p>
<p><strong>C. The leader who thrives will look different from the leader who thrived in the last decade.</strong> Not “more soft skills” as a slogan. A different job: <strong>business outcomes with technical teeth</strong>, not people-process as the product.</p>
<p>Those three reinforce each other. When coding throughput jumps, the bottleneck moves to judgment — what to build, what to kill, what interface to freeze, what risk is acceptable. Judgment is leadership work. Ticket wrangling is not.</p>
<h2>What the surviving leader actually owns</h2>
<p>The leaders I would hire into a flatter org share a pattern:</p>
<ul>
<li><strong>They own business outcomes, not just a people inventory.</strong> Headcount is an input. Revenue, reliability, adoption, and cost of delay are the scoreboard.</li>
<li><strong>They understand product, engineering, operations, and how sales actually sells.</strong> Enough to argue tradeoffs in the same language as partners — not as a translator who only ships Jira.</li>
<li><strong>They think in systems, not only sprints.</strong> Feedback loops, failure domains, platform interfaces, incentive design. Sprint hygiene still matters; it is not the strategy.</li>
<li><strong>They keep enough technical depth that returning to IC tomorrow is credible.</strong> Not day-one expert on every stack — but not a spectator in design reviews either.</li>
</ul>
<p>The old clean split — pure management track vs pure IC track — is already leaking. Hybrid value is the default. Specialization remains; <strong>pure process ownership without domain force</strong> does not.</p>
<h2>If I stepped into an EM role tomorrow</h2>
<p>I would do three things immediately, and keep doing them:</p>
<ol>
<li><strong>Stay technically current.</strong> Read designs. Ship small things. Touch the tools agents use. You cannot review what you cannot feel.</li>
<li><strong>Invest hard in strategic and business thinking.</strong> How the company makes money, where margin dies, which customers actually drive the roadmap. Engineering leadership without that is interior decoration.</li>
<li><strong>Do not sit on laurels.</strong> Title and past scope depreciate faster when the production function changes. Relevance is a practice, not a plaque.</li>
</ol>
<p>None of that is glamorous. All of it is insurance against becoming the middle layer the spreadsheet deletes first.</p>
<h2>The org shape that follows</h2>
<p>I expect organizations that work under AI-heavy delivery to look roughly like this:</p>
<ul>
<li><strong>Executives at the top owning strategy</strong> — fewer, clearer bets.</li>
<li><strong>A lean middle layer owning specific business goals</strong> — not generic “people management” for its own sake.</li>
<li><strong>Highly productive ICs (and agent fleets) supported by real tooling</strong> — platforms, evals, review culture, operational ownership.</li>
</ul>
<p>That is dramatically flatter than the multi-layer manager stacks many companies still run. Flatter is not kinder by default. It is <strong>less room for ambiguity about who owns the outcome.</strong> If your role only exists to pass information up and down, the system will route around you.</p>
<h2>The question that matters</h2>
<p>The question is not whether this shift is coming. Parts of it are already priced into how aggressive companies staff management. The question is whether you are preparing: technical depth, business literacy, and systems judgment — or optimizing for a job description that is evaporating.</p>
<p>What are you doing right now to stay relevant as the middle of the org chart thins?</p></content><category term="Blog"/><category term="engineering-leadership"/><category term="management"/><category term="product"/></entry><entry><title>Near-real-time is a product promise, not a Kafka cluster</title><link href="https://www.mohitranka.com/blog/near-real-time-is-a-product-promise/" rel="alternate"/><published>2026-07-16T10:00:00+05:30</published><updated>2026-07-16T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2026-07-16:/blog/near-real-time-is-a-product-promise/</id><summary type="html"><p>“Near-real-time” is often used as a synonym for “we bought a streaming stack.” On GTM data platforms, that confusion is expensive. The product promise is about <strong>whether a decision-maker can trust a number in time</strong> — not whether an event log is busy. At LinkedIn, the Enterprise Data Platform (EDP) was …</p></summary><content type="html"><p>“Near-real-time” is often used as a synonym for “we bought a streaming stack.” On GTM data platforms, that confusion is expensive. The product promise is about <strong>whether a decision-maker can trust a number in time</strong> — not whether an event log is busy. At LinkedIn, the Enterprise Data Platform (EDP) was intended as the governed center for GTM datasets. BI teams on Power BI and Tableau still ran on Hadoop-era batch paths. They already had data. What they did not have was a reason to treat EDP as the path that made GTM analytics true, timely, and durable as legacy pipelines were marked for deprecation. That gap is a product-promise gap.</p>
<figure class="post-figure">
<img src="https://www.mohitranka.com/images/blog/near-real-time-is-a-product-promise.jpg" alt="Abstract illustration of a clock and data streams feeding a product dashboard" loading="lazy" width="1200" height="675">
<figcaption>
<span class="fig-caption">Near-real-time is a promise about the age of truth a consumer can act on—not a badge for a streaming cluster.</span>
</figcaption>
</figure>
<!--more-->
<h2>What consumers actually bought</h2>
<p>When sales, marketing, or BI stakeholders ask for better data timing, they rarely mean “please introduce a new consumer group.” They mean things like:</p>
<ul>
<li>The dashboard does not contradict the operational truth for long.</li>
<li>A GTM metric used in a weekly motion is not secretly a stale extract.</li>
<li>There is one place to stand when two numbers disagree.</li>
</ul>
<p>Those are promises to <strong>named consumers</strong>. For us, the critical early consumers included BI engineering and the GTM partners who lived in their dashboards. If those consumers do not move, the platform’s internal latency graphs are cosplay.</p>
<h2>A definition that forces honesty</h2>
<p>I use a boring definition: for dataset <em>D</em> and consumer <em>C</em>, the promise is that the <strong>age of usable data</strong> at the moment <em>C</em> acts stays inside an agreed bound — under normal conditions — with a clear story when it does not. Unpack it in platform language:</p>
<ul>
<li><strong>Named dataset</strong> — a sales or GTM entity teams recognize, not “the lake.”</li>
<li><strong>Named consumer</strong> — Power BI workbook owners, Tableau extracts, revenue analytics — not “downstream.”</li>
<li><strong>Usable</strong> — passes the governance and correctness bar, not merely “row arrived.”</li>
<li><strong>Agreed bound</strong> — may start as a milestone (“on EDP before deprecation date, with acceptable query latency”) before it becomes a polished SLO.</li>
<li><strong>Failure story</strong> — what the business does when the path is wrong: freeze a pipeline, pin a version, staff a war room — not “we’ll check the dashboard Monday.”</li>
</ul>
<p>If you cannot name <em>D</em> and <em>C</em>, you are not ready to sell near-real-time.</p>
<h2>Why “we can already get the data” kills platform promises</h2>
<p>The BI objection was rational: existing Hadoop-based jobs still produced outputs. From their seat, migration was risk without immediate upside. So the competing product was not another vendor. It was <strong>the legacy batch path that still worked</strong>. Near-real-time platforms lose to “good enough yesterday” until:</p>
<ol>
<li>the old path has an end date,</li>
<li>the new path is staffed for adopters,</li>
<li>query performance after cutover is somebody’s on-call problem.</li>
</ol>
<p>We treated those as part of the promise design. EDP engineers helped migrate. Connectors reduced friction into Power BI and Tableau. Deprecation deadlines made dual-running finite. Performance work made “governed” not mean “slower.” Streaming could have been part of some paths. It was not the adoption strategy.</p>
<h2>Hidden decisions that are actually the product</h2>
<h3>Where is the system of record?</h3>
<p>If EDP and a legacy pipeline disagree, which number is allowed to win in a QBR deck? Until that is explicit, faster pipelines just produce faster arguments.</p>
<h3>Is the promise contractual or best-effort?</h3>
<p>Deprecating Hadoop paths is a contractual move: the company is choosing a continuity posture. Best-effort “please try EDP” will lose to local convenience forever.</p>
<h3>Who pays for fan-out?</h3>
<p>Every BI team and every GTM dataset multiplies support surface. Platforms that treat every new consumer as free eventually stop being able to keep any promise.</p>
<h3>What happens during catch-up and cutover?</h3>
<p>Migrations create windows where two truths coexist. The product promise must cover the dual-run period — or you will invent tribal knowledge in Slack.</p>
<h2>The shape that worked: phases, not a big-bang bus</h2>
<p>The useful architecture picture was organizational as much as technical:</p>
<ol>
<li><strong>Executive buy-in</strong> — GTM data consolidation as strategy, not a side project. </li>
<li><strong>MVP on high-value sales datasets</strong> — prove the promise where pain is visible. </li>
<li><strong>Expand across sales and revenue consumers</strong> — widen the interface only after the path works. </li>
<li><strong>Broader GTM standardization</strong> — reduce snowflake pipelines. </li>
<li><strong>Governance and optimization</strong> — make the default path the boring path.</li>
</ol>
<p>That sequence is how you keep a freshness/correctness promise while the org is still learning to trust the platform. “Put everything on a stream in quarter one” is usually a way to buy complexity before you have consumers.</p>
<h2>Questions I ask before approving “let’s go real-time”</h2>
<ul>
<li>Which decision improves if this dataset gets younger — and who makes that decision?</li>
<li>What is the current path’s delay, really (including BI extracts and cache TTLs)?</li>
<li>What dies when we succeed — the legacy job, or only our spare time?</li>
<li>Who is staffed to migrate the first consumers?</li>
<li>What query latency is acceptable in the tools people actually use?</li>
<li>What do we measure weekly that a VP would recognize?</li>
</ul>
<p>If the answers are all technology choices, the promise is not ready.</p>
<h2>Closing</h2>
<p>Near-real-time is not a badge for an architecture review. It is a <strong>promise about time, truth, and behavior when truth is late</strong>. EDP’s lesson was blunt: the company does not get that promise when a platform exists. It gets that promise when critical consumers — here, BI on GTM data — run on the governed path, and the old batch defaults are allowed to end. Build streams when they earn their keep. Build the consumer promise first.</p></content><category term="Blog"/><category term="data-platforms"/><category term="distributed-systems"/><category term="reliability"/></entry><entry><title>The platform team’s real job is interfaces</title><link href="https://www.mohitranka.com/blog/platform-teams-real-job-is-interfaces/" rel="alternate"/><published>2025-11-13T10:00:00+05:30</published><updated>2025-11-13T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2025-11-13:/blog/platform-teams-real-job-is-interfaces/</id><summary type="html"><p>At LinkedIn, the Enterprise Data Platform (EDP) was meant to be the centralized way GTM teams managed and consumed datasets. On paper, that is a clear platform charter. In practice, a platform is only real when its <strong>interfaces get adopted</strong> — including by teams that already have a path that “works …</p></summary><content type="html"><p>At LinkedIn, the Enterprise Data Platform (EDP) was meant to be the centralized way GTM teams managed and consumed datasets. On paper, that is a clear platform charter. In practice, a platform is only real when its <strong>interfaces get adopted</strong> — including by teams that already have a path that “works.” The hard case was BI. Power BI and Tableau teams were still living on Hadoop-based pipelines scheduled for deprecation. They could already get data. EDP was strategically important and still optional in their week. Without them, EDP could not become the source of truth for GTM analytics no matter how good the internals looked. That is an interface problem, not a cluster problem.</p>
<figure class="post-figure">
<img src="https://www.mohitranka.com/images/blog/platform-teams-real-job-is-interfaces.jpg" alt="Abstract illustration of connecting building blocks and interface ports between teams" loading="lazy" width="1200" height="675">
<figcaption>
<span class="fig-caption">Platforms earn leverage when the interfaces other teams stand on are clear, adoptable, and owned.</span>
</figcaption>
</figure>
<!--more-->
<h2>The interface was “how BI gets governed data”</h2>
<p>When platform teams say interface, they often mean API shape or event schema. Here the consumer-facing interface was broader:</p>
<ul>
<li>How does a BI engineer get a trusted dataset into a dashboard workflow?</li>
<li>Who pays the migration cost?</li>
<li>What happens to the old path, and when?</li>
<li>Is query performance acceptable the week after cutover?</li>
</ul>
<p>EDP’s storage format and pipeline elegance did not answer those questions. Until they were answered, BI had no reason to reorder priorities. <strong>Lesson:</strong> product teams do not consume your architecture diagrams. They consume time-to-success, predictability, and whether the platform team shows up when the path is rocky.</p>
<h2>Adoption failed as an org problem first</h2>
<p>This was easy to misread as stubbornness. It was incentives.</p>
<ul>
<li>BI teams were measured on analytics delivery, not on platform migration.</li>
<li>Legacy pipelines still produced outputs.</li>
<li>Migration looked like unfunded work with downside risk (latency, rework, surprise breakage).</li>
<li>EDP’s long-term governance story was real — and still abstract compared to this quarter’s dashboards.</li>
</ul>
<p>Engineering alignment inside the data org was necessary and insufficient. Without a business framing, “please adopt EDP” is a favor request.</p>
<h2>Executive sponsorship changed the type of conversation</h2>
<p>I partnered with senior leaders on the data and BI side so EDP adoption was not a side quest. The useful reframe was not “modernize for us.” It was:</p>
<ul>
<li>Hadoop paths were going away.</li>
<li>Continuing to depend on them was a <strong>continuity risk</strong>, not a neutral default.</li>
<li>EDP was the consolidation path for governed GTM analytics — not a nice-to-have alternate store.</li>
</ul>
<p>That shift moved the discussion from technical preference to operating risk. Platforms that cannot get that sentence said out loud usually stall in permanent pilot mode.</p>
<h2>We lowered the price of yes</h2>
<p>Even with sponsorship, asking BI to self-fund a migration would have failed slowly. The decision that mattered operationally: <strong>EDP engineers would take migration load</strong> — connectors, pairing, performance work — not only publish docs and wish for pull requests. Concretely, that meant:</p>
<ul>
<li>Pre-built paths into Power BI and Tableau so “get data from EDP” was not a research project.</li>
<li>Shared work on query performance so cutover did not mean slower dashboards.</li>
<li>A migration toolkit and recurring office hours to kill blockers in public, early.</li>
<li>Alignment with sales and marketing stakeholders who depended on the outputs, not only the producers.</li>
</ul>
<p>This is platform-as-product without the theater: the job-to-be-done was “keep GTM analytics working on a governed foundation,” and we priced the platform to make that job rational.</p>
<h2>Deprecation is part of the interface</h2>
<p>Enabling a new path without disabling the old one is how companies collect platforms. The adoption plan included the unglamorous half:</p>
<ul>
<li>Clear deprecation milestones for legacy Hadoop pipelines.</li>
<li>Time-bound dual running where needed.</li>
<li>A definition of done that included <strong>turning things off</strong>, not only turning EDP on.</li>
</ul>
<p>Governance is not a slide about ownership. Governance is whether the abandoned path still quietly feeds production dashboards six months later.</p>
<h2>What changed</h2>
<p>Within a concentrated push — on the order of a quarter for the BI migration motion — Power BI and Tableau usage moved onto EDP for the scoped GTM paths we targeted. Redundant pipeline surface area could be decommissioned. Freshness and governance improved because fewer competing “sources of truth” were allowed to linger. I care less about the trophy phrasing than about the mechanism: <strong>sponsorship + funded migration + forced deprecation + BI-shaped interfaces.</strong></p>
<h2>Where EMs earn their keep on platform teams</h2>
<p>The technical work was real. The EM job showed up in places that never appear in a system design doc:</p>
<ul>
<li>Spending team capacity on someone else’s migration so the company’s data model could converge.</li>
<li>Holding a deprecation line when temporary extensions would have been easier.</li>
<li>Making sure performance issues after cutover were our problem, not a gotcha that punished adopters.</li>
<li>Translating platform risk into language executives will prioritize.</li>
</ul>
<p>Special cases still existed — BI tools always have them — but they were absorbed into connectors and support rituals, not infinite private forks.</p>
<h2>Closing</h2>
<p>Platform teams love infrastructure. Infrastructure is not the job. The job is to define, evolve, and defend <strong>interfaces other teams can build on</strong> — including the incentives, migration labor, and deprecation schedule that make those interfaces real. EDP did not become the GTM source of truth when it launched. It became the source of truth when BI could succeed on it, and the old paths were allowed to die.</p></content><category term="Blog"/><category term="engineering-leadership"/><category term="platforms"/><category term="developer-tooling"/></entry><entry><title>When I still choose a relational database</title><link href="https://www.mohitranka.com/blog/when-i-still-choose-a-relational-database/" rel="alternate"/><published>2025-03-13T10:00:00+05:30</published><updated>2025-03-13T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2025-03-13:/blog/when-i-still-choose-a-relational-database/</id><summary type="html"><p>In 2013 I wrote <a href="https://www.mohitranka.com/blog/rdbms-vs-nosql/">RDBMS vs. NOSQL?</a> as a pushback against fashion. The fashion changed costumes — document stores, wide-column, NewSQL, "Postgres is fine," "everything in the lakehouse" — but the underlying mistake did not: <strong>picking a datastore from a blog post instead of from access patterns and failure modes.</strong></p>
<figure class="post-figure">
<img src="https://www.mohitranka.com/images/blog/when-i-still-choose-a-relational-database.jpg" alt="Illustration of an ordered data table foundation beside scattered documents" loading="lazy" width="1200" height="675">
<figcaption>
<span class="fig-caption">A relational …</span></figcaption></figure></summary><content type="html"><p>In 2013 I wrote <a href="https://www.mohitranka.com/blog/rdbms-vs-nosql/">RDBMS vs. NOSQL?</a> as a pushback against fashion. The fashion changed costumes — document stores, wide-column, NewSQL, "Postgres is fine," "everything in the lakehouse" — but the underlying mistake did not: <strong>picking a datastore from a blog post instead of from access patterns and failure modes.</strong></p>
<figure class="post-figure">
<img src="https://www.mohitranka.com/images/blog/when-i-still-choose-a-relational-database.jpg" alt="Illustration of an ordered data table foundation beside scattered documents" loading="lazy" width="1200" height="675">
<figcaption>
<span class="fig-caption">A relational spine is often the boring default until access patterns and failure modes force a different shape.</span>
</figcaption>
</figure>
<!--more-->
<p>This is the sequel I would write to myself: when I still choose a relational database in 2026, and when I do not.</p>
<h2>The default remains boring on purpose</h2>
<p>My default for a new product backend is still a managed relational database (usually PostgreSQL) with:</p>
<ul>
<li>Clear schema ownership</li>
<li>Migrations as code</li>
<li>Backups and point-in-time recovery you have actually restored</li>
<li>Connection pooling and boring observability</li>
<li>A plan for read scale that is not "wishful replicas"</li>
</ul>
<p>Why? Because most products are <strong>transactional workflows with relationships</strong>: users, accounts, permissions, orders, configurations, audit trails. Relational databases are extraordinarily good at that shape. They also match how humans reason about correctness. If your primary problems are multi-row invariants, ad hoc query flexibility, and operational maturity, starting elsewhere is often self-inflicted difficulty.</p>
<h2>When relational is the right call</h2>
<p>I lean relational when several of these are true:</p>
<p><strong>1. Strong invariants matter more than infinite write scale.</strong><br>
Money, entitlements, identity bindings, "exactly one active X per Y." If the business invariant is relational, fighting the model is expensive.</p>
<p><strong>2. Query patterns are still evolving.</strong><br>
Early products change questions weekly. A well-modeled schema with indexes beats a write-optimized store that makes new questions painful.</p>
<p><strong>3. The team’s operational muscle is SQL-shaped.</strong><br>
A perfect paper architecture with zero operators is worse than a known system with runbooks. Skills are part of architecture.</p>
<p><strong>4. Multi-entity transactions simplify the product.</strong><br>
Saga forests can be correct. They are also a tax. If a single-node or lightly clustered RDBMS can hold the transactional core, keep the core small and sharp.</p>
<p><strong>5. You need ecosystem gravity.</strong><br>
ORMs, migration tools, BI access, hiring, incident folklore — relational ecosystems are deep. That is not marketing; it is time-to-recovery.</p>
<h2>When relational becomes the wrong center</h2>
<p>I move work <em>out</em> of the primary OLTP database when:</p>
<p><strong>1. Write throughput or state size exceeds honest vertical + read-replica plans.</strong><br>
Not vanity metrics — measured saturation, vacuum pain, replica lag that product feels, backup windows that scare you.</p>
<p><strong>2. Access patterns are append-heavy and query-narrow.</strong><br>
Event logs, high-volume telemetry, massive multi-tenant time series. Different stores exist for reasons.</p>
<p><strong>3. Fan-out reads need specialized shapes.</strong><br>
Search, graph traversal at scale, feature stores, geospatial at high QPS — often better as <strong>derived systems</strong> fed from a transactional core.</p>
<p><strong>4. Availability topology requirements outgrow your relational operator model.</strong><br>
Some multi-region active-active stories are possible with modern relational systems; many are still research projects wearing production clothes. Be honest about the topology you can run.</p>
<p><strong>5. Schema flexibility is real, not aesthetic.</strong><br>
Truly heterogeneous documents with little shared query structure can be a poor fit. "We might need flexibility later" is not a requirement.</p>
<h2>The architecture that aged best for me</h2>
<p>The pattern I trust:</p>
<ol>
<li><strong>Transactional system of record</strong> in relational (or something with equal invariant strength).</li>
<li><strong>Explicit integration events</strong> when other systems must react — versioned, owned, documented.</li>
<li><strong>Derived read models</strong> for specialized query paths.</li>
<li><strong>Analytical systems</strong> for heavy aggregation — do not abuse OLTP as a warehouse.</li>
<li><strong>Clear rules for dual writes</strong> — preferably avoid them; if not, make correctness visible.</li>
</ol>
<p>The relational database is not the universe. It is often the <strong>spine</strong>.</p>
<h2>A related lesson from platform service boundaries</h2>
<p>Not every “monolith vs microservices” fight is a datastore fight — but the same judgment applies. On an EDP self-serve portal, we had a month-long freeze over whether new UI workflows had to live entirely in a monolithic backend or split early for independent iteration. The useful answer was hybrid: keep <strong>core platform capabilities</strong> where integration and invariants dominate; split <strong>fast-changing workflows</strong> where deploy independence pays for the seam cost; put <strong>metadata</strong> in a system designed for lifecycle, not in accidental dual writes. That is the same muscle as choosing a relational spine: <strong>optimize for ownership and correctness at the center, derive or split at the edges, refuse fashion-driven rewrites that stop delivery.</strong></p>
<h2>What changed since 2013</h2>
<p>A few updates to my older self:</p>
<ul>
<li><strong>Managed Postgres (and peers) are much better.</strong> Failover, backups, and scaling knobs improved. That raises the bar for "we need NoSQL to be reliable."</li>
<li><strong>JSON columns done carefully</strong> cover many former document-store arguments — without abandoning transactions.</li>
<li><strong>CDC became mainstream.</strong> Turning relational truth into streams is a product pattern, not a science fair.</li>
<li><strong>Distributed SQL exists</strong> and can be right — but it is still a specialist choice with specialist costs.</li>
<li><strong>Lakehouses ate a chunk of analytics</strong>, which is good: stop pretending your OLTP DB is a lake.</li>
</ul>
<p>What did <em>not</em> change: <strong>NoSQL is not a personality.</strong> It is a set of tradeoffs around consistency, query power, operational complexity, and scale.</p>
<h2>Decision checklist I actually use</h2>
<p>Before approving a non-relational primary store:</p>
<ol>
<li>Which queries and invariants are first-class in v1?</li>
<li>What is the 12-month data size and QPS with ugly margins?</li>
<li>What does multi-row correctness look like, and who enforces it?</li>
<li>How do we migrate schema or access patterns six months in?</li>
<li>Who is on-call, and what have they operated before?</li>
<li>Can we explain the choice in one paragraph without vendor slogans?</li>
</ol>
<p>If step 6 requires a conference talk, the choice may still be right — but it needs more proof.</p>
<h2>Closing</h2>
<p>I still choose a relational database when I want <strong>correctness, evolvable queries, and operational boredom</strong> around the core of a product. I choose other systems when the access pattern or scale curve makes that boredom impossible. The winning move is rarely "pick a side." It is <strong>keep a sharp system of record, derive aggressively, and refuse to let fashion rename your requirements.</strong> If you want the older, more argumentative version of this stance, it is still here: <a href="https://www.mohitranka.com/blog/rdbms-vs-nosql/">RDBMS vs. NOSQL?</a>. The industry moved. The need for judgment did not.</p></content><category term="Blog"/><category term="data-platforms"/><category term="databases"/><category term="architecture"/></entry><entry><title>Freshness SLOs: the metric product teams actually feel</title><link href="https://www.mohitranka.com/blog/freshness-slos-the-metric-product-teams-feel/" rel="alternate"/><published>2024-07-11T10:00:00+05:30</published><updated>2024-07-11T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2024-07-11:/blog/freshness-slos-the-metric-product-teams-feel/</id><summary type="html"><p>API latency SLOs changed how we run services. GTM data work taught me the sibling idea the hard way: product teams do not experience your job-success chart. They experience a dashboard that is late, wrong, or disagrees with another “official” number. At LinkedIn, BI teams on Power BI and Tableau …</p></summary><content type="html"><p>API latency SLOs changed how we run services. GTM data work taught me the sibling idea the hard way: product teams do not experience your job-success chart. They experience a dashboard that is late, wrong, or disagrees with another “official” number. At LinkedIn, BI teams on Power BI and Tableau could already pull data from Hadoop-era batch paths while the Enterprise Data Platform (EDP) was supposed to become the governed center for GTM datasets. Pipelines can be “green” while the business still feels stale or fragmented truth. That feeling is a freshness and trust problem — even when nobody is paging on consumer lag.</p>
<figure class="post-figure">
<img src="https://www.mohitranka.com/images/blog/freshness-slos-the-metric-product-teams-feel.jpg" alt="Illustration of a freshness gauge on an operations dashboard" loading="lazy" width="1200" height="675">
<figcaption>
<span class="fig-caption">Product teams feel freshness as trust in the number—not as job-success charts on a platform dashboard.</span>
</figcaption>
</figure>
<!--more-->
<h2>Latency is not freshness</h2>
<p>A fast dashboard query can still serve an extract that is hours old. A successful batch job can still leave sales and marketing acting on last night’s world. Latency asks: how long did this request take?<br>
Freshness asks: <strong>how old is the truth this decision used?</strong> If you only measure the serving API, you will celebrate the wrong layer.</p>
<h2>Define freshness in the consumer’s language</h2>
<p>For EDP-shaped work, a useful definition was: for a named GTM dataset and a named BI consumer, how old is the data at the moment someone uses it in a dashboard decision — and is it the governed path? That forces specifics:</p>
<ul>
<li><strong>Dataset</strong> — a sales or GTM entity people recognize.</li>
<li><strong>Consumer</strong> — a Power BI or Tableau workflow, not “analytics.”</li>
<li><strong>Usable</strong> — on the platform you claim is source of truth, not a leftover pipeline.</li>
<li><strong>Age</strong> — including batch boundaries and BI-side refresh behavior, not only platform ingest time.</li>
</ul>
<p>You can formalize that into an SLO later. First you need a sentence a BI lead and a platform EM would both underline.</p>
<h2>Promises that change behavior (not charts for their own sake)</h2>
<p>On the BI migration, the promises that changed behavior were not abstract percentiles on a poster. They looked like:</p>
<ul>
<li><strong>Continuity:</strong> legacy Hadoop paths had deprecation dates — staying put was a risk posture, not a neutral default.</li>
<li><strong>Funded cutover:</strong> EDP engineers helped migrate; adopters were not asked to donate a quarter of calendar alone.</li>
<li><strong>Acceptable query performance after switch</strong> — so “governed” did not mean “slower dashboards.”</li>
<li><strong>Finite dual-running</strong> — two truths were a migration window, not a lifestyle.</li>
</ul>
<p>Call those SLOs if your org has the discipline. Call them <strong>operating promises</strong> if you are still early. Either way, something has to hurt when the promise breaks — attention, prioritization, or the ability to keep the old path alive. I will not pretend every dataset had a polished error budget. What we had was executive visibility, migration milestones, and a definition of done that included turning legacy paths off.</p>
<h2>What to measure along a real path</h2>
<p>For a BI consumer, time and trust hide in stages:</p>
<ol>
<li>Source systems and upstream delay </li>
<li>Platform ingest and validation on EDP </li>
<li>Dataset readiness / governance checks </li>
<li>BI tool refresh or extract behavior </li>
<li>Cache and workbook-level assumptions</li>
</ol>
<p>Teams often discover the villain is not the fanciest processor. It is an extract schedule, a dual pipeline, or a cutover that never finished. Instrument the path your consumers actually use.</p>
<h2>Correctness sits beside freshness</h2>
<p>Fresh wrong data is worse than slightly stale right data. During migration, dual sources create a special failure mode: two numbers, both “recent,” different owners. Correctness signals that mattered in practice:</p>
<ul>
<li>Is this dataset on the governed path? </li>
<li>Are legacy feeds still quietly serving production workbooks? </li>
<li>Do critical fields null out or drift after cutover? </li>
<li>Can someone name the system of record when a QBR fights itself?</li>
</ul>
<p>Freshness without governance optimizes for speed of confusion.</p>
<h2>How we introduced the promise for BI</h2>
<p>The sequence that worked was not “roll out SLO framework company-wide”:</p>
<ol>
<li>Pick the consumer class already on the critical path (BI for GTM). </li>
<li>Make adoption a leadership priority with a continuity narrative. </li>
<li>Staff the migration so the new path is cheaper than it looks. </li>
<li>Put dates on deprecation. </li>
<li>Fix performance issues as platform bugs, not user error. </li>
<li>Only then talk about tightening time bounds dataset by dataset.</li>
</ol>
<p>If you start with twenty datasets and a perfect taxonomy, you will get a wiki. If you start with one embarrassing consumer journey, you might get a habit.</p>
<h2>The objection we actually heard</h2>
<p><strong>“We can already get the data.”</strong><br>
Yes. That is why platform-only arguments fail. Answer with end-of-life for the old path, labor for the new one, and proof that cutover does not degrade the dashboard.</p>
<p><strong>“Batch is fine.”</strong><br>
Sometimes it is. Then write a batch promise (“available by time T for the weekly motion”) and still own correctness and deprecation. Batch without a promise is how shadow pipelines live forever.</p>
<h2>Closing</h2>
<p>Product teams feel freshness as trust in the number. Platforms earn that trust when named consumers run on a governed path, old paths can die, and someone is accountable when the age of truth is wrong. Whether you brand it SLO or operating promise matters less than whether the company can point to <strong>dataset + consumer + bound + owner</strong> — and whether missing it changes next week’s work.</p></content><category term="Blog"/><category term="data-platforms"/><category term="reliability"/><category term="observability"/></entry><entry><title>Saying no as a platform EM without becoming the villain</title><link href="https://www.mohitranka.com/blog/saying-no-as-a-platform-em/" rel="alternate"/><published>2023-11-08T10:00:00+05:30</published><updated>2023-11-08T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2023-11-08:/blog/saying-no-as-a-platform-em/</id><summary type="html"><p>Platform engineering managers do not run out of good ideas. They run out of capacity to say yes to every reasonable request without wrecking the shared system. The skill is not blunt refusal. It is <strong>no with a path</strong> — specific enough that partners can actually execute, firm enough that your …</p></summary><content type="html"><p>Platform engineering managers do not run out of good ideas. They run out of capacity to say yes to every reasonable request without wrecking the shared system. The skill is not blunt refusal. It is <strong>no with a path</strong> — specific enough that partners can actually execute, firm enough that your team is not a free consulting desk. Two nos from EDP work at LinkedIn taught me more than any generic prioritization framework.</p>
<figure class="post-figure">
<img src="https://www.mohitranka.com/images/blog/saying-no-as-a-platform-em.jpg" alt="Illustration of a fork in the road with one path closed and an alternate route open" loading="lazy" width="1200" height="675">
<figcaption>
<span class="fig-caption">A useful platform “no” names the constraint and funds a path—not a silent block.</span>
</figcaption>
</figure>
<!--more-->
<h2>Why platform “no” feels personal</h2>
<p>Product and BI partners are graded on what ships this quarter. Platform teams are graded on leverage, reliability, and whether the company still has one coherent data model next year. So when you decline an unfunded migration or a rewrite dressed up as a preference, it can sound like you do not care. Sometimes that critique is fair. Often the real problem is that you said no without offering a way through. A villain blocks silently. A partner names the constraint and funds a path.</p>
<h2>Principles I actually use</h2>
<ol>
<li><strong>Company throughput beats local speed.</strong> A one-week special case that creates a permanent support branch is not kindness.</li>
<li><strong>A delayed honest yes beats a fake yes.</strong> A Jira ticket with no staffing is a lie.</li>
<li><strong>Tradeoffs go in writing the same day.</strong> Memory is political; notes are kinder.</li>
</ol>
<p>Everything below is those three principles in concrete form. If you cannot show how a decision shows up in ownership, metrics, and day-to-day work, it will not survive the next roadmap fight.</p>
<h2>Case A — No to “just adopt the platform”</h2>
<p><strong>Request (implied):</strong> BI teams on Power BI and Tableau should move to EDP because it is the strategic GTM data platform. <strong>Reality:</strong> They could already get data from Hadoop-based pipelines. Migration looked like unfunded risk. And EDP could not become the source of truth without them. <strong>The no:</strong> No to a pure mandate without labor — “adopt EDP” as a favor to the platform team. <strong>The path:</strong></p>
<ul>
<li>Executive sponsorship that framed legacy pipeline deprecation as <strong>continuity risk</strong>, not taste.</li>
<li><strong>EDP engineers assigned to migration work</strong>, not only documentation.</li>
<li>Connectors into Power BI and Tableau.</li>
<li>Query performance work so cutover did not punish adopters.</li>
<li>Tooling, office hours, and <strong>hard deprecation milestones</strong> so dual-running actually ended.</li>
</ul>
<p>That package is a no to magical adoption and a yes to an expensive but real interface change. Within a concentrated push — about a quarter for the BI motion we scoped — the migration stuck, and legacy surface area could shrink. <strong>Pattern:</strong> If the consumer has no incentive, your “no” is to unfunded asks; your “yes” is to change incentives and put a price on the work.</p>
<h2>Case B — No to both pure extremes in an architecture fight</h2>
<p><strong>Request (implied):</strong> Pick monolith <em>or</em> microservices for an EDP self-serve portal — each side sure the other choice was malpractice. <strong>Reality:</strong> The disagreement went public, ownership collapsed, and delivery froze for roughly a month. <strong>The no:</strong> No to a binary holy war. No to indefinite debate. No to “loudest critique wins.” <strong>The path:</strong></p>
<ul>
<li>Structured design review with <strong>written criteria</strong> (scalability, maintainability, speed, ownership).</li>
<li><strong>Time-boxed POCs</strong> from both approaches instead of slide wars.</li>
<li>A neutral senior engineer in the room.</li>
<li>Explicit coaching on ownership and influence in 1:1s — not only technical arbitration.</li>
<li>A <strong>hybrid decision</strong>: core platform capabilities stayed integrated with the EDP backend; more dynamic portal workflows could be separate services; metadata lifecycle centralized (we used DataHub) rather than reinvented.</li>
</ul>
<p>Execution resumed about a week after the decision landed. The portal itself took months, with real adoption along the way. <strong>Pattern:</strong> Sometimes the EM’s no is to false dichotomies. The funded path is a hybrid with proofs, not a victory lap for one camp.</p>
<h2>Make yes expensive in the right way</h2>
<p>Temporary exceptions will exist — dual pipelines during migration, transitional architecture branches. Price them:</p>
<ul>
<li><strong>Time-bounded</strong> (a deprecation date, not vibes)</li>
<li><strong>Owned</strong> (a named team for breakage)</li>
<li><strong>Visible</strong> (on a list leadership can see)</li>
<li><strong>Removable</strong> (exit criteria written down)</li>
</ul>
<p>Free, quiet exceptions are how platforms drown.</p>
<h2>Roadmaps are how you say no at scale</h2>
<p>One-off negotiation does not survive GTM scope. The EDP sales/GTM program needed an explicit sequence:</p>
<ol>
<li>Buy-in and prioritization </li>
<li>MVP on high-value sales datasets </li>
<li>Expand across sales/revenue consumers </li>
<li>Broader GTM standardization </li>
<li>Governance and optimization</li>
</ol>
<p>That roadmap is a machine for “not yet.” Without it, every dataset is an emergency, and every emergency becomes a yes.</p>
<h2>Protect the team without hiding behind them</h2>
<p>“The team is busy” is weak if you cannot show the math. Better:</p>
<ul>
<li>Here is committed platform work (migration staffing, deprecation, portal seams).</li>
<li>Here is what we will not staff this quarter.</li>
<li>Here is the escalation if the business wants to reorder.</li>
</ul>
<p>Take heat in partner forums so individual engineers are not negotiating company priority alone.</p>
<h2>Closing</h2>
<p>Saying no as a platform EM is stewardship of shared constraints. On EDP, the nos that mattered were: <strong>no unfunded adoption</strong>, and <strong>no architecture theater that freezes delivery</strong>. The yeses were expensive on purpose — engineers on migration, connectors, deprecation, POCs, hybrid seams. If partners can see the path, you are not the villain. You are how the company keeps one platform instead of twelve.</p></content><category term="Blog"/><category term="engineering-leadership"/><category term="platforms"/><category term="management"/></entry><entry><title>Identity systems fail socially before they fail cryptographically</title><link href="https://www.mohitranka.com/blog/identity-systems-fail-socially/" rel="alternate"/><published>2023-03-08T10:00:00+05:30</published><updated>2023-03-08T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2023-03-08:/blog/identity-systems-fail-socially/</id><summary type="html"><p>When people talk about identity and SSO, they reach for algorithms: token lifetimes, key rotation, SAML vs OIDC, session fixation. Those details matter. In systems I have built and operated, the outages and near-misses that hurt most started earlier — as <strong>social and product failures</strong> wearing security clothing.</p>
<figure class="post-figure">
<img src="https://www.mohitranka.com/images/blog/identity-systems-fail-socially.jpg" alt="Illustration of keys, badges, and people connected in a trust network" loading="lazy" width="1200" height="675">
<figcaption>
<span class="fig-caption">Identity systems often …</span></figcaption></figure></summary><content type="html"><p>When people talk about identity and SSO, they reach for algorithms: token lifetimes, key rotation, SAML vs OIDC, session fixation. Those details matter. In systems I have built and operated, the outages and near-misses that hurt most started earlier — as <strong>social and product failures</strong> wearing security clothing.</p>
<figure class="post-figure">
<img src="https://www.mohitranka.com/images/blog/identity-systems-fail-socially.jpg" alt="Illustration of keys, badges, and people connected in a trust network" loading="lazy" width="1200" height="675">
<figcaption>
<span class="fig-caption">Identity systems often fail first as ownership and session-semantics problems, not as crypto bugs.</span>
</figcaption>
</figure>
<!--more-->
<h2>Identity is a dependency graph of humans</h2>
<p>An identity platform is not only a service. It is:</p>
<ul>
<li>Who is allowed to grant access </li>
<li>How contractors, partners, and acquisitions show up </li>
<li>What "logout" means across devices and apps </li>
<li>Which team gets paged when login breaks on a Sunday </li>
<li>How quickly a leaver loses access in reality, not in policy PDFs</li>
</ul>
<p>Crypto bugs are rare relative to <strong>misowned workflows</strong>. The system can be textbook-correct and still fail the organization.</p>
<h2>Failure mode 1: Ambiguous system of record for "who is this person?"</h2>
<p>Enterprises collect identities the way rivers collect silt: HRIS, directories, partner IdPs, legacy user tables, support tools that mint exceptions. If you cannot answer "what is the canonical identifier, and who can change it?" you will eventually:</p>
<ul>
<li>Duplicate humans </li>
<li>Orphan entitlements </li>
<li>Merge the wrong accounts </li>
<li>Build reconciliation jobs that become the real product</li>
</ul>
<p>SSO does not fix identity entropy. It multiplies whatever model you already have. <strong>Design move:</strong> pick a primary subject key strategy, document merge/split procedures, and make account recovery a first-class flow — not a Zendesk folklore.</p>
<h2>Failure mode 2: Special cases without expiry</h2>
<p>Sales needs a demo tenant. Support needs impersonation. A partner needs a long-lived integration user. Security accepts a temporary bypass. Temporary becomes permanent. Permanent becomes unmonitored. Unmonitored becomes the breach path. <strong>Design move:</strong> every exception has an owner, an expiry, metrics, and a removal plan. Impersonation and break-glass are products with audit trails, not Slack approvals.</p>
<h2>Failure mode 3: Logout and session mental models differ by app</h2>
<p>Users think "I logged out." Your distributed sessions think "two refresh tokens and a cache entry are still valid." Mobile thinks something else. A partner app never got the memo. This is not merely UX. It is an access-control bug with a friendly face. <strong>Design move:</strong> define session lifecycle as an explicit cross-app contract. Test logout like you test login. Include shared devices and support scenarios.</p>
<h2>Failure mode 4: Rollouts that treat auth like a feature flag toy</h2>
<p>Identity changes have asymmetric risk. A broken profile color is annoying. A broken token validation is company-wide stoppage. Yet teams still ship auth changes like UI tweaks: wide rollouts, thin dashboards, no rehearsal of rollback. <strong>Design move:</strong> progressive exposure, synthetic login journeys per IdP, clear rollback that does not require a hero, and change freezes around known peak login events.</p>
<h2>Failure mode 5: Ownership is "security and also platform and also the app"</h2>
<p>When login fails, three teams page each other. When it works, nobody funds hardening. Diffused ownership produces brittle reliability. <strong>Design move:</strong> a single operational owner for the login path, with written dependencies on IdP, DNS, email, device services, and app session layers. Security sets policy; platform runs the path; apps integrate against a stable interface. Blurry RACI is an availability risk.</p>
<h2>What good looks like</h2>
<p>Strong identity programs I have seen share a few traits:</p>
<ol>
<li><strong>Boring standards on the outside</strong> (OIDC/SAML done plainly) </li>
<li><strong>Strict internal models</strong> for subjects, credentials, devices, and grants </li>
<li><strong>Auditability as a product feature</strong> </li>
<li><strong>Recovery and leaver flows tested</strong>, not assumed </li>
<li><strong>Customer-visible status</strong> when auth is degraded </li>
<li><strong>Load and failure testing of login</strong>, not only of core APIs</li>
</ol>
<p>Notice how little of that is "pick the trendy token format."</p>
<h2>Questions before you add another identity feature</h2>
<ul>
<li>Who is the human-level source of truth? </li>
<li>What is the break-glass path, and who audits it? </li>
<li>How does a user understand their sessions? </li>
<li>What happens to downstream caches on revoke? </li>
<li>Which team’s error budget does login reliability consume? </li>
<li>Can we re-run last quarter’s incidents as game days?</li>
</ul>
<p>If those answers are soft, new federation features will add surface area, not safety.</p>
<h2>Closing</h2>
<p>Identity systems do fail cryptographically — and you should hire people who care about that deeply. But if you only harden tokens while leaving ownership, exceptions, session semantics, and rollout discipline vague, you will still fail. They fail socially first: unclear truth, unowned edges, temporary forever, and teams that cannot coordinate under stress. Build the social protocol as carefully as the crypto protocol. Users feel both. Attackers only need one to be weak.</p></content><category term="Blog"/><category term="identity"/><category term="security"/><category term="distributed-systems"/></entry><entry><title>What “led the web launch” taught me about constraints</title><link href="https://www.mohitranka.com/blog/web-launch-constraints/" rel="alternate"/><published>2022-07-06T10:00:00+05:30</published><updated>2022-07-06T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2022-07-06:/blog/web-launch-constraints/</id><summary type="html"><p>At Postman, “put the product on the web” was not a greenfield rewrite. It was a constraint problem: a desktop-native API tool used by millions of developers, enterprise pressure for browser access, browser security that blocked the old execution model, and a conference date that would not move. I was …</p></summary><content type="html"><p>At Postman, “put the product on the web” was not a greenfield rewrite. It was a constraint problem: a desktop-native API tool used by millions of developers, enterprise pressure for browser access, browser security that blocked the old execution model, and a conference date that would not move. I was the engineering manager accountable for cross-functional delivery — architecture choices, security and infra dependencies, product scope, and keeping the team focused when the path was still uncertain. What follows is what that launch actually taught me.</p>
<figure class="post-figure">
<img src="https://www.mohitranka.com/images/blog/web-launch-constraints.jpg" alt="Illustration of browser and desktop windows bridged together" loading="lazy" width="1200" height="675">
<figcaption>
<span class="fig-caption">Launching a trusted desktop product on the web is a constraint problem: security, parity, and an immovable date.</span>
</figcaption>
</figure>
<!--more-->
<h2>We were borrowing trust, not inventing it</h2>
<p>Postman already had a reputation. Developers had muscle memory for collections, environments, and the desktop workflow. A weak web surface would not be judged as “v1 of a new product.” It would be judged as Postman getting worse. That constraint changed prioritization:</p>
<ul>
<li>Core workflow parity beat architectural purity.</li>
<li>Predictable behavior beat clever browser tricks.</li>
<li>Explicit “not on web yet” beat silent missing features.</li>
</ul>
<p>Trust is spendable once. We treated every launch-day gap as a brand risk, not a backlog curiosity.</p>
<h2>The existing system was a stakeholder</h2>
<p>Greenfield essays assume you choose the stack. We inherited runtimes, sync assumptions, offline collaboration expectations, and a large surface area of API tooling behavior. The desktop app was not legacy to be embarrassed about — it was the system of record for how users worked. The hard question was never “can we draw a web architecture?” It was:</p>
<ul>
<li>What must be reused so results stay correct?</li>
<li>What must be isolated so the browser can ship?</li>
<li>What bugs will the web amplify because usage patterns change?</li>
</ul>
<p>Treating the existing product as a stakeholder forced interface thinking: web was a new client of a product system, not a parallel fantasy product.</p>
<h2>Browser security forced a hybrid execution model</h2>
<p>Desktop Postman could talk to the network like a normal app. Browsers cannot. CORS, sandboxing, and the lack of unrestricted local network access were not edge cases — they were the product. We ended up with multiple execution paths, each owning a real constraint:</p>
<ul>
<li><strong>Browser Agent</strong> — run requests directly when the browser is allowed to.</li>
<li><strong>Cloud Agent</strong> — execute in a cloud-hosted environment when the browser cannot reach the target cleanly (cross-origin and related limits).</li>
<li><strong>Desktop Agent</strong> — bridge the web UI to a local/on-prem network when the user’s world is not reachable from the public cloud.</li>
</ul>
<p>That hybrid model was the architectural heart of the launch. It was also an organizational heart: security, infra, and product had to agree on what “send request” meant in three different trust domains. If there is one technical lesson I would keep from the project, it is this: <strong>when the environment cannot support your old runtime assumptions, make the execution model explicit.</strong> Hiding three behaviors behind one button without a design is how you get support chaos.</p>
<h2>Launch day is a reliability event</h2>
<p>Postman on the Web was announced at POSTCON. Missing the date was not a soft failure mode. That does not mean we shipped fantasy scope. It means readiness was defined as:</p>
<ul>
<li>a user journey that worked under real constraints,</li>
<li>a rollout plan that could expand,</li>
<li>and a team that knew what was deliberately later.</li>
</ul>
<p>We phased capability instead of pretending the first public cut was the end state: start with constrained access patterns, then enable richer API execution, with the cloud execution path continuing to mature after the headline launch. Marketing owns the keynote. Engineering owns the degradation and expansion story. A launch is not a timestamp. It is a reliability event with an audience.</p>
<h2>Cross-team coordination was the critical path</h2>
<p>The longest pole was rarely a single function. Security, infrastructure, performance, and product had legitimate, conflicting optimization targets. In that environment, “the engineers will figure it out in Slack” is not a plan. What worked in practice:</p>
<ul>
<li><strong>Executive air cover for dedicated capacity</strong> — without it, every dependency team optimizes for their prior roadmap.</li>
<li><strong>A weekly cross-functional sync</strong> whose job was unblocking, not status theater.</li>
<li><strong>Written scope decisions</strong> — what was in for conference day, what was explicit debt, who owned the follow-through.</li>
</ul>
<p>My calendar was part of the architecture. Ambiguity multiplies under deadline pressure; decision logs shrink it.</p>
<h2>Performance was a product constraint, not polish</h2>
<p>API collections can be huge. A desktop WebView habit does not automatically become a good browser experience. Large histories and large collections will punish naive rendering. We invested in boring, necessary work: more efficient history/state handling, lazy loading, virtualized UI for large lists. That work is easy to dismiss as polish until a power user loads a real workspace and the tab melts. Developer products have an unforgiving feedback loop. Users can tell when the runtime is lying, when the UI is papering over cost, and when error messages are decorative. They will also write about it publicly. That is part of the market.</p>
<h2>What I would repeat</h2>
<ol>
<li><strong>Write non-goals as carefully as goals.</strong> Conference-day success needs a spine, not a vision deck.</li>
<li><strong>Instrument journeys, not only services.</strong> “Request failed” is incomplete without <em>which agent path</em> and <em>which constraint</em>.</li>
<li><strong>Rehearse partial failure</strong> — auth issues, agent unavailability, dependency brownouts — not only happy-path demos.</li>
<li><strong>Staff the week after launch like it is part of launch.</strong> The real traffic pattern arrives after the keynote.</li>
<li><strong>Protect engineers from thrash</strong> by batching stakeholder input; panic multiplies bad architectural shortcuts.</li>
</ol>
<h2>What I would avoid</h2>
<ul>
<li>Betting the public launch on an unfinished platform rewrite that is “almost ready.”</li>
<li>Hiding scope cuts inside the word “polish.”</li>
<li>Success metrics only a marketing team can love.</li>
<li>Hero culture that makes the second week impossible to staff.</li>
</ul>
<h2>Impact, carefully stated</h2>
<p>We hit the conference launch window. The web surface became a real product path, not a demo — with cloud execution continuing to land on its own schedule. Adoption afterward made the strategic point obvious: users wanted Postman without installing a desktop app first, and the company was no longer only a local-first tool. Exact figures belong in contexts where they can be sourced and defended. The leadership lesson does not depend on a screenshot of a dashboard: <strong>the hybrid execution model plus phased delivery was the only way to respect browser constraints without abandoning the desktop product’s trust.</strong></p>
<h2>Closing</h2>
<p>“Led the web launch” sounds like a milestone. The work was constraint management: borrowed trust, inherited systems, browser security, conference time, cross-team conflict, and performance under real collections. Code expressed the answers. The answers were the constraints we were willing to name early — and the execution model we built so users did not have to understand them all at once.</p></content><category term="Blog"/><category term="engineering-leadership"/><category term="product"/><category term="developer-tooling"/></entry><entry><title>Incubating 0→1 beside a mature product</title><link href="https://www.mohitranka.com/blog/postman-labs-0-to-1/" rel="alternate"/><published>2021-06-15T10:00:00+05:30</published><updated>2021-06-15T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2021-06-15:/blog/postman-labs-0-to-1/</id><summary type="html"><p>When I was at Postman, the core product was already the default API client for a huge HTTP/HTTPS world — on the order of tens of millions of developers on desktop. That success created a sharp problem: <strong>how do you explore what comes after HTTP without slowing the product everyone …</strong></p></summary><content type="html"><p>When I was at Postman, the core product was already the default API client for a huge HTTP/HTTPS world — on the order of tens of millions of developers on desktop. That success created a sharp problem: <strong>how do you explore what comes after HTTP without slowing the product everyone already depends on?</strong> Postman Labs was our answer: a small unit with a charter to incubate 0→1 work — new protocols and paradigms — without turning every experiment into a core-roadmap hostage situation.</p>
<figure class="post-figure">
<img src="https://www.mohitranka.com/images/blog/postman-labs-0-to-1.jpg" alt="Illustration of a small lab greenhouse beside a solid product building" loading="lazy" width="1200" height="675">
<figcaption>
<span class="fig-caption">0→1 incubation works when it sits beside a mature product with a real graduation path—not as endless side quests.</span>
</figcaption>
</figure>
<!--more-->
<h2>The challenge: the market moved past “REST in a GUI”</h2>
<p>Developers were increasingly living with:</p>
<ul>
<li><strong>WebSockets</strong> for persistent, bidirectional sessions </li>
<li><strong>gRPC</strong> for efficient, schema-driven service APIs </li>
<li><strong>GraphQL</strong> and other non-CRUD shapes of interface</li>
</ul>
<p>Internally we had ideas. What we lacked was a structured way to <strong>validate, build, and kill or graduate</strong> them while the core team stayed focused on the desktop product’s quality and scale. Putting every bet into the main feature factory would have meant either:</p>
<ul>
<li>starving core reliability and UX, or </li>
<li>shipping “innovation” at the speed of a mature backlog.</li>
</ul>
<p>Labs existed to refuse that false choice.</p>
<h2>Structure for speed (and containment)</h2>
<p>We did not invent another feature team with the same process tax as core. Labs was intentionally different:</p>
<ol>
<li><strong>Smaller team</strong> drawn from engineers who already knew Postman’s product DNA. </li>
<li><strong>Less process theater</strong> — lean experiments instead of full SDLC cosplay for every spike. </li>
<li><strong>One North Star metric per initiative</strong> — a single definition of success so debates stayed grounded. </li>
<li><strong>Explicit separation</strong> from core delivery so a failed experiment did not become a multi-quarter core commitment by accident.</li>
</ol>
<p>Charter in two axes:</p>
<ul>
<li><strong>Breadth</strong> — protocols and shapes beyond HTTP. </li>
<li><strong>Depth</strong> — personas and workflows (testing, automation, CI-shaped use) that the HTTP client alone did not own.</li>
</ul>
<p>Independence was not isolation from users. It was isolation from the wrong kind of backlog pressure.</p>
<h2>A three-phase model that de-risked 0→1</h2>
<p>Every Labs initiative had to earn the next phase:</p>
<h3>Phase 1 — Feasibility and market validation</h3>
<p>Customer conversations, lightweight proofs, honest “who hurts without this?” If the answer was only “it would be cool,” it did not proceed.</p>
<h3>Phase 2 — MVP and dogfooding</h3>
<p>Internal builds used by Postman’s own engineers. Usability and correctness issues showed up before a public audience.</p>
<h3>Phase 3 — Limited beta and public validation</h3>
<p>Constrained rollout, measure adoption and engagement, then decide: graduate into the main product, iterate, or stop.</p>
<p>That sequence sounds obvious. The discipline is stopping between phases. Core roadmaps often skip phase 1 and call a half-built feature a launch.</p>
<h2>WebSockets: persistent sessions are not “requests with extra steps”</h2>
<p>WebSockets were an early Labs bet because real-time systems (chat, feeds, IoT-style control planes, live tooling) do not fit the request/response mental model Postman had optimized for. <strong>Hard parts:</strong></p>
<ul>
<li>Long-lived bidirectional connections instead of discrete calls </li>
<li>Auth flows that must hold for a session, not only a single hit </li>
<li>UX for streams of events over time, not one response panel</li>
</ul>
<p><strong>What we built toward:</strong> a first-class WebSocket client experience — connect, send/receive, inspect event history, support practical auth patterns. <strong>Outcome:</strong> WebSockets did not stay a lab toy; support graduated into Postman’s broader API development surface. That graduation path was the point of Labs.</p>
<h2>gRPC: schemas and streams in a product trained on text HTTP</h2>
<p>gRPC was growing fast in backend-heavy environments. Postman’s muscle memory was text-centric HTTP. gRPC forced different questions:</p>
<ul>
<li>How do users work with <strong>Protobuf</strong> contracts inside the product? </li>
<li>How do unary and streaming calls show up in a composer UX? </li>
<li>How do serialization mistakes become debuggable instead of opaque binary pain?</li>
</ul>
<p><strong>Execution themes:</strong> schema-aware composition, streaming-aware request/response handling, and making “what did I just send?” inspectable for developers who live in Postman daily. <strong>Outcome:</strong> gRPC testing/debugging became a real product capability in the same family as REST workflows — not a separate science project users had to leave Postman for.</p>
<h2>What scaled beyond the first bets</h2>
<p>When early graduates worked, Labs stopped being only a temporary squad. The incubation pattern — validate, dogfood, beta, graduate — became a reusable company muscle. Later product bets (including automation-shaped work such as Flows-class ideas) could reuse the same organizational shape: <strong>explore beside core, then merge what earns users.</strong> Eventually Labs-shaped work attracted clearer funding and leadership attention. That is the healthy end state: not a permanent rebel base, but a proven path for 0→1 inside a company that also has to protect a mature product.</p>
<h2>What I would repeat as an EM</h2>
<ol>
<li><strong>Separate exploration capacity from core SLA capacity</strong> — or core always wins and innovation becomes slideware. </li>
<li><strong>One success metric per bet</strong> — multi-metric dashboards hide kill decisions. </li>
<li><strong>Dogfood before marketing</strong> — especially for developer tools; your engineers are harsh, useful users. </li>
<li><strong>Graduation criteria in writing</strong> — “done in Labs” must mean something operationally. </li>
<li><strong>Protect the team from identity crisis</strong> — Labs is not “the people who do side quests”; it is a product strategy role.</li>
</ol>
<h2>What I would avoid</h2>
<ul>
<li>Innovation theater with no kill switch </li>
<li>Hiding Labs work so core is surprised at graduation </li>
<li>Measuring success only by launches, not by retained use </li>
<li>Staffing Labs only with people core “can spare” forever</li>
</ul>
<h2>Closing</h2>
<p>Incubating 0→1 beside a mature product is a leadership design problem: process, incentives, and graduation rules — not only prototype velocity. Postman Labs worked when it had a <strong>narrow charter</strong>, <strong>phased validation</strong>, and a path for WebSockets, gRPC, and similar bets to become real product surfaces without forcing the entire company to pretend it was still a startup with nothing to lose. If your core product is already loved, that is not a reason to stop exploring. It is a reason to explore <strong>on purpose</strong>.</p></content><category term="Blog"/><category term="engineering-leadership"/><category term="product"/><category term="developer-tooling"/></entry><entry><title>How I review a distributed design in 45 minutes</title><link href="https://www.mohitranka.com/blog/how-i-review-a-distributed-design/" rel="alternate"/><published>2021-03-03T10:00:00+05:30</published><updated>2021-03-03T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2021-03-03:/blog/how-i-review-a-distributed-design/</id><summary type="html"><p>The most expensive design reviews I have run were not missing a box on a diagram. They were missing a decision. One of them froze delivery on an EDP self-serve portal for about a month while two strong engineers disagreed in public about monolith versus microservices — and ownership quietly collapsed …</p></summary><content type="html"><p>The most expensive design reviews I have run were not missing a box on a diagram. They were missing a decision. One of them froze delivery on an EDP self-serve portal for about a month while two strong engineers disagreed in public about monolith versus microservices — and ownership quietly collapsed. This is how I run a distributed design review when time is short, using that conflict as the worked example. The forty-five minutes are a filter for <strong>danger and indecision</strong>, not a substitute for deep design.</p>
<figure class="post-figure">
<img src="https://www.mohitranka.com/images/blog/how-i-review-a-distributed-design.jpg" alt="Illustration of an architecture whiteboard with boxes, arrows, and coffee cups" loading="lazy" width="1200" height="675">
<figcaption>
<span class="fig-caption">A short design review is a filter for clear promises, truth models, and decisions—not a theater of diagrams.</span>
</figcaption>
</figure>
<!--more-->
<h2>The situation the review had to unstick</h2>
<p>We needed a self-serve portal on top of the Enterprise Data Platform: dataset registration, governance controls, access workflows, metadata — so teams could manage lifecycle without filing tickets into oblivion. Two credible positions formed:</p>
<ul>
<li><strong>Stay close to the monolithic EDP backend</strong> — simpler integration, less duplication, faster delivery on shared infra.</li>
<li><strong>Split into microservices early</strong> — independent iteration on portal features without waiting on the monolith.</li>
</ul>
<p>Both sides had technical merit. The failure mode was social: critique moved into open forums in a way that undermined the owner, the owner disengaged, and execution stopped. Stakeholders saw silence. That is a design-process failure, not only an architecture debate.</p>
<h2>Minutes 0–5: What user promise are we keeping?</h2>
<p>Before hexagons, I want one paragraph:</p>
<ul>
<li>Who uses the portal?</li>
<li>What can they do without a human intermediary?</li>
<li>What is explicitly out of scope for v1?</li>
</ul>
<p>For us: GTM/data producers and consumers managing dataset lifecycle — registration, access, metadata — not “rebuild EDP as microservices.” If the promise is fuzzy, stop. Architecture will invent scope.</p>
<h2>Minutes 5–15: Where does truth live?</h2>
<p>I care about systems of record more than service count. Questions that mattered on the portal:</p>
<ul>
<li>Which actions must be consistent with core EDP backend behavior on day one?</li>
<li>Which workflows change weekly and need independent deploy cadence?</li>
<li>Where does dataset metadata live so the portal is not a second brain?</li>
</ul>
<p>We eventually used a hybrid truth model: <strong>core platform capabilities stayed integrated with the existing EDP backend</strong>; <strong>more dynamic workflows</strong> (access requests, tagging-style features) could stand as separate services; <strong>metadata</strong> was centralized with a system fit for dataset lifecycle (in our case, Apache DataHub) so the portal was not inventing yet another catalog. Red flags in any review:</p>
<ul>
<li>“Both systems will stay in sync” with no mechanism.</li>
<li>Every feature forced into one deployability story.</li>
<li>Metadata treated as a UI detail.</li>
</ul>
<h2>Minutes 15–25: What fails—technically and organizationally?</h2>
<p>Classic distributed questions still apply: timeouts, dual writes, partial deploy, replay. On this project the binding failure was different:</p>
<ul>
<li>What happens if the owning engineer stops driving?</li>
<li>What happens if disagreement becomes a public referendum every week?</li>
<li>What happens if leadership hears only one side’s framing?</li>
</ul>
<p>A design review that ignores ownership and decision rights will produce a beautiful diagram and a still project. I schedule the technical argument <strong>inside a structured forum</strong> with criteria — not in drive-by threads. If you need blame-free space, create it deliberately.</p>
<h2>Minutes 25–35: How will we choose without infinite debate?</h2>
<p>Opinion without evidence burns weeks. The intervention that worked:</p>
<ol>
<li><strong>Write evaluation criteria</strong> before picking a winner: scalability, maintainability, delivery speed, long-term ownership.</li>
<li><strong>Force small proofs</strong> — both approaches get a time-boxed spike/POC against the criteria.</li>
<li><strong>Bring a neutral senior engineer</strong> into the room to pressure-test both sides without owning either ego.</li>
<li><strong>Time-box the decision</strong> — the review ends with a path, not a sequel meeting.</li>
</ol>
<p>This is operability of the <em>decision</em>, not only of the service.</p>
<h2>Minutes 35–40: People side (do not skip)</h2>
<p>Distributed design is done by humans with status and career goals. In parallel with the technical path:</p>
<ul>
<li>Rebuild ownership with the engineer who had stepped back — silence is not an acceptable escalation strategy.</li>
<li>Coach the critic on influence: staff-level impact includes <em>how</em> you challenge, not only that you are right.</li>
<li>Make expectations explicit in 1:1s so the project is not a proxy war.</li>
</ul>
<p>Skip this and the hybrid architecture will still die in the next disagreement.</p>
<h2>Minutes 40–45: Decision and conditions</h2>
<p>We did not pick a pure monolith or a pure microservice rewrite. We picked a <strong>hybrid</strong>:</p>
<ul>
<li>Core platform features (registration, governance controls tightly bound to EDP) stayed where integration cost dominated.</li>
<li>Dynamic portal features that needed independent iteration moved toward separate services.</li>
<li>Metadata lifecycle was centralized rather than re-implemented.</li>
</ul>
<p>Approve with conditions, in writing: what is in v1, what is explicitly later, who owns the seams, when the next review is if assumptions fail. Verbal “sounds good” evaporates. Written conditions survive contact with calendars.</p>
<h2>The diagram I want if I only get one</h2>
<p>Trade three layered architecture posters for either:</p>
<ul>
<li>a <strong>sequence</strong> of one user action through registration → metadata → access, or </li>
<li>a <strong>side-by-side POC scorecard</strong> against the agreed criteria.</li>
</ul>
<p>Sequence diagrams and scorecards reveal lies that box diagrams hide.</p>
<h2>Anti-patterns this freeze taught me</h2>
<ul>
<li>Public architecture criticism that bypasses the owner.</li>
<li>Binary holy wars (monolith vs microservices) without workload specifics.</li>
<li>Design review as spectator sport for leadership without a decision owner.</li>
<li>EM waiting too long to facilitate because “they’re seniors, they’ll figure it out.”</li>
<li>Approving to end discomfort rather than risk.</li>
</ul>
<h2>What unblocked looked like</h2>
<p>Once criteria, POCs, and a hybrid decision landed, the freeze broke quickly — on the order of a week to resume real execution — and the portal shipped on a timeline measured in months with meaningful adoption. The architecture mattered. The restored ownership mattered more.</p>
<h2>Closing</h2>
<p>In forty-five minutes you will not finish a distributed design. You can learn whether the team has a clear promise, a coherent truth model, a way to decide, and a human ownership path. On platform work, the last item is not soft. It is how delivery fails first.</p></content><category term="Blog"/><category term="distributed-systems"/><category term="engineering-leadership"/><category term="architecture"/></entry><entry><title>Introducing on-call without burning the team</title><link href="https://www.mohitranka.com/blog/introducing-on-call-without-burnout/" rel="alternate"/><published>2020-11-10T10:00:00+05:30</published><updated>2020-11-10T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2020-11-10:/blog/introducing-on-call-without-burnout/</id><summary type="html"><p>At Postman, my team owned a large-scale platform surface used by millions of developers. What we did not own — formally — was a <strong>predictable operational response</strong>. Production issues and public GitHub noise were handled ad hoc. Someone jumped in, or everyone hesitated. Retrospectives lacked a clear accountable role. The system worked …</p></summary><content type="html"><p>At Postman, my team owned a large-scale platform surface used by millions of developers. What we did not own — formally — was a <strong>predictable operational response</strong>. Production issues and public GitHub noise were handled ad hoc. Someone jumped in, or everyone hesitated. Retrospectives lacked a clear accountable role. The system worked until it did not, and then it worked by heroics. We needed on-call. We also needed engineers who still wanted to build product after the rotation.</p>
<figure class="post-figure">
<img src="https://www.mohitranka.com/images/blog/introducing-on-call-without-burnout.jpg" alt="Illustration of a calm night operations desk with a pager and schedule board" loading="lazy" width="1200" height="675">
<figcaption>
<span class="fig-caption">On-call is a product and staffing design: clear ownership without making heroics the default.</span>
</figcaption>
</figure>
<!--more-->
<h2>What was broken before process</h2>
<p>Without a rotation, four things piled up:</p>
<ol>
<li><strong>Unclear ownership</strong> — incidents waited on “who feels responsible today.” </li>
<li><strong>Context-switch tax</strong> — feature work and firefighting shared the same brains without boundaries. </li>
<li><strong>Uneven load</strong> — the same people always raised their hands. </li>
<li><strong>Weak learning loops</strong> — retros had symptoms, not a role that carried fixes week to week.</li>
</ol>
<p>For a developer-facing platform, user-visible breakage is not a side channel. It is the product. Treating ops as optional was a product decision, whether we admitted it or not.</p>
<h2>Design goals</h2>
<p>I wanted a system that was:</p>
<ul>
<li><strong>Explicit</strong> — someone is primary, always. </li>
<li><strong>Fair</strong> — load rotates; it does not stick to volunteers. </li>
<li><strong>Bounded</strong> — on-call is a job for a window, not a personality type. </li>
<li><strong>Educational</strong> — the whole team sees production, not only a martyr subset. </li>
<li><strong>Humane</strong> — no permanent page-from-bed culture dressed up as commitment.</li>
</ul>
<p>Reliability that depends on burnout is just deferred attrition.</p>
<h2>The model we ran</h2>
<h3>Primary and secondary</h3>
<ul>
<li><strong>Primary</strong> — dedicated to monitoring, incident response, and triage (including GitHub-facing noise). Feature delivery expectations drop for that window on purpose. </li>
<li><strong>Secondary</strong> — backup for major incidents; not a stealth second primary for every ping.</li>
</ul>
<p>If primary is still expected to hit the same sprint commitments, you do not have on-call. You have theater.</p>
<h3>Two-week shifts with a handoff pattern</h3>
<p>Engineers rotated on a <strong>two-week</strong> cadence. A common pattern was primary one window, then secondary the next — so knowledge transferred and no one lived forever in the blast radius. Back-to-back primary stretches were treated as a smell, not a badge.</p>
<h3>Handoff as a team ritual</h3>
<p>Weekly team time included an <strong>on-call handoff</strong>, not only standup status. Primary walked through:</p>
<ul>
<li>Incidents and resolutions </li>
<li>Adjacent system issues that might hit us next </li>
<li>Notable GitHub / support themes </li>
<li>Follow-ups that needed owners beyond the shift</li>
</ul>
<p>That ritual turned private pager pain into shared product knowledge.</p>
<h3>Retros and playbooks</h3>
<p>Major incidents got structured retros. Recurring issues earned <strong>playbooks</strong> — step-by-step paths so the next primary was not rediscovering folklore at 1 a.m. Playbooks are how on-call becomes a team asset instead of tribal knowledge in one engineer’s head.</p>
<h3>Psychological safety and load management</h3>
<ul>
<li>No expectation of endless consecutive primaries. </li>
<li>Balance with feature work across the quarter so people are not typed as “ops only.” </li>
<li>Managers (me included) treated page load and fairness as staffing concerns, not only engineer grit.</li>
</ul>
<h2>What improved</h2>
<p>With clear ownership:</p>
<ul>
<li><strong>Faster response</strong> — we saw on the order of a <strong>~40% reduction in incident response time</strong> once roles were unambiguous (directionally; treat it as an order-of-magnitude win from process, not a lab result). </li>
<li><strong>More systematic triage</strong> — fewer “is anyone looking at this?” gaps. </li>
<li><strong>Better morale</strong> — predictable pain beats random pain. </li>
<li><strong>Broader operational skill</strong> — more engineers touched production reality. </li>
<li><strong>Earlier fixes</strong> — primaries had space to chip at known sharp edges before they became SEVs.</li>
</ul>
<p>Stability improved because response became a designed system, not a personality contest.</p>
<h2>What I would tell another EM before day one</h2>
<ol>
<li><strong>Write who is primary in a place the team actually looks.</strong> </li>
<li><strong>Cut feature load for primary</strong> or you will train people to ignore the pager. </li>
<li><strong>Ship handoff and playbooks in the first month</strong>, not after the third outage. </li>
<li><strong>Measure response and fairness</strong>, not only uptime. </li>
<li><strong>Defend the rotation against “just this once” exceptions</strong> from leadership — exceptions are how volunteers reappear.</li>
</ol>
<h2>Failure modes to avoid</h2>
<ul>
<li>On-call as punishment for the least political engineers </li>
<li>Secondary as free extra primary </li>
<li>Retros without owners or due dates </li>
<li>Alert noise so high that everyone mutes everything </li>
<li>Celebrating heroes instead of fixing the systems that required them</li>
</ul>
<h2>Closing</h2>
<p>Introducing on-call is not a tooling purchase. It is a <strong>product and staffing decision</strong>: user trust requires an accountable human path, and that path must be sustainable. At Postman, primary/secondary roles, two-week rotations, handoffs, retros, and playbooks turned operational ownership from ad hoc heroics into something the team could carry — and still ship. If your platform is already large and your response is still “whoever notices,” you do not have a reliability gap only. You have a leadership design gap. Close it on purpose.</p></content><category term="Blog"/><category term="reliability"/><category term="engineering-leadership"/><category term="platforms"/></entry><entry><title>GTM datasets need data contracts</title><link href="https://www.mohitranka.com/blog/gtm-datasets-need-data-contracts/" rel="alternate"/><published>2020-07-01T10:00:00+05:30</published><updated>2020-07-01T10:00:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2020-07-01:/blog/gtm-datasets-need-data-contracts/</id><summary type="html"><p>When a go-to-market dashboard is wrong, nobody says “the warehouse is eventually consistent.” They say the number is wrong — and they stop trusting the platform. On LinkedIn’s Enterprise Data Platform (EDP) work, the failure mode was rarely a missing chart type. It was <strong>informal truth</strong>: datasets without clear producers …</p></summary><content type="html"><p>When a go-to-market dashboard is wrong, nobody says “the warehouse is eventually consistent.” They say the number is wrong — and they stop trusting the platform. On LinkedIn’s Enterprise Data Platform (EDP) work, the failure mode was rarely a missing chart type. It was <strong>informal truth</strong>: datasets without clear producers, consumers, freshness expectations, or a path that BI could rely on when Hadoop-era pipelines still “worked.” That is a data-contract problem, whether or not you use the word contract.</p>
<figure class="post-figure">
<img src="https://www.mohitranka.com/images/blog/gtm-datasets-need-data-contracts.jpg" alt="Illustration of two parties exchanging a contract over data folders" loading="lazy" width="1200" height="675">
<figcaption>
<span class="fig-caption">GTM datasets become trustworthy when producers and consumers share explicit contracts—not informal folklore.</span>
</figcaption>
</figure>
<!--more-->
<h2>The contract is the product boundary</h2>
<p>A data contract is a written agreement between people who produce a dataset and people who depend on it:</p>
<ul>
<li>What the dataset means (grain, keys, critical fields)</li>
<li>Who owns changes</li>
<li>How fresh and complete it must be for its main consumers</li>
<li>Who may use it, and for what</li>
<li>What happens when the shape changes</li>
<li>Which path is the system of record when two feeds disagree</li>
</ul>
<p>Without that, you have tables and jobs — not a product interface. EDP’s strategic job was to become the governed center for GTM data. BI teams on Power BI and Tableau did not move because a platform existed; they moved when the <strong>interface of getting trustworthy data</strong> became clearer, cheaper, and eventually mandatory as legacy paths aged out.</p>
<h2>Which GTM datasets need contracts (almost all that matter)</h2>
<p><strong>Definitely:</strong></p>
<ul>
<li>Pipeline and revenue metrics that show up in leadership reviews </li>
<li>Account, lead, and opportunity-shaped datasets used across tools </li>
<li>Any feed BI materializes into workbooks that drive weekly motions </li>
<li>Datasets used for access decisions, eligibility, or customer-facing ops </li>
<li>Shared “golden” entities multiple teams join in different ways</li>
</ul>
<p><strong>Lighter-weight is fine for:</strong></p>
<ul>
<li>Truly exploratory sandboxes with no production consumers </li>
<li>One-off extracts with an explicit expiry</li>
</ul>
<p>If a number can start an argument in a QBR, it deserves a contract.</p>
<h2>What we needed in practice (not a 40-page template)</h2>
<p>Keep contracts short enough that producers and BI partners will actually read them:</p>
<ol>
<li><strong>Dataset name</strong> and owning team </li>
<li><strong>Grain</strong> (what one row means) </li>
<li><strong>Critical fields</strong> and allowed null behavior </li>
<li><strong>Primary consumers</strong> (e.g. Power BI / Tableau paths, sales analytics) </li>
<li><strong>Freshness / readiness expectation</strong> — even if it starts as “available on EDP before legacy deprecation,” not a perfect percentile </li>
<li><strong>Change policy</strong> — notice, versioning, who approves breaking changes </li>
<li><strong>System of record</strong> during dual-run periods </li>
<li><strong>Support path</strong> — where breakages go (not a random Slack thread)</li>
</ol>
<p>On EDP, connectors, migration staffing, and deprecation dates were how contracts became real. A wiki table alone does not change incentives.</p>
<h2>Why platform adoption without contracts fails</h2>
<p>BI’s rational objection was: “We can already get the data.” Informal sources always feel free until:</p>
<ul>
<li>Two dashboards disagree </li>
<li>A legacy pipeline is turned off </li>
<li>A field changes meaning and nobody tells the workbook owner </li>
<li>Query performance after cutover becomes “the platform’s problem” with no owner</li>
</ul>
<p>Contracts force those conversations <strong>before</strong> the incident. They also make deprecation fair: you cannot retire a path nobody documented as non-authoritative.</p>
<h2>Evaluation is part of the contract</h2>
<p>Do not separate “data quality” from “dashboard quality.” Ship and migration gates should include:</p>
<ul>
<li>Consumer path checks (does the BI workflow still resolve?) </li>
<li>Row-count / null-rate sanity on critical fields </li>
<li>Explicit dual-run comparisons while both paths live </li>
<li>A named human who can freeze a bad publish</li>
</ul>
<p>When the contract breaks, something visible should fail before the QBR does.</p>
<h2>Organizational pattern that worked</h2>
<ul>
<li><strong>Producers</strong> own correctness and change communication </li>
<li><strong>Platform (EDP)</strong> owns enforcement, discovery, access patterns, and migration leverage </li>
<li><strong>BI / analytics partners</strong> own consumer semantics and workbook impact </li>
<li><strong>Leadership</strong> owns deprecation as continuity policy, not a style preference</li>
</ul>
<p>Shared Slack channels are not a substitute for ownership. Funded migration and executive sponsorship were how EDP contracts left the slide deck.</p>
<h2>A sequence I recommend for new GTM datasets</h2>
<ol>
<li>Write the decision the dataset supports in one sentence. </li>
<li>Name the first production consumer (often a BI path). </li>
<li>Draft the contract <em>before</em> scaling access. </li>
<li>Put the dataset on the governed platform path. </li>
<li>Dual-run only with an end date. </li>
<li>Turn off the informal path.</li>
</ol>
<p>Teams love to start at “expose the table.” Steps 1–3 are where trust is designed.</p>
<h2>Closing</h2>
<p>EDP did not earn “source of truth” status by existing. It earned it when GTM consumers — especially BI — could depend on <strong>named datasets with owners, expectations, and an end to competing pipelines</strong>. Call that a data contract, a product interface, or an operating promise. Just do not ship GTM data as folklore and hope governance appears later.</p></content><category term="Blog"/><category term="data-platforms"/><category term="platforms"/><category term="product"/></entry><entry><title>RDBMS vs. NOSQL?</title><link href="https://www.mohitranka.com/blog/rdbms-vs-nosql/" rel="alternate"/><published>2013-06-29T00:16:00+05:30</published><updated>2013-06-29T00:16:00+05:30</updated><author><name>Mohit Ranka</name></author><id>tag:www.mohitranka.com,2013-06-29:/blog/rdbms-vs-nosql/</id><summary type="html"><blockquote><p>From our own experience designing and operating a highly available, highly scalable ecommerce platform, we have come to realize that relational databases should only be used when an application really needs the complex query, table join and transaction capabilities of a full-blown relational database. In all other cases, when such …</p></blockquote></summary><content type="html"><blockquote><p>From our own experience designing and operating a highly available, highly scalable ecommerce platform, we have come to realize that relational databases should only be used when an application really needs the complex query, table join and transaction capabilities of a full-blown relational database. In all other cases, when such relational features are not needed, a NoSQL database service like DynamoDB offers a simpler, more available, more scalable and ultimately a lower cost solution.</p></blockquote>
<pre><code> — Werner Vogels, CTO Amazon.com on when to use RDBMS
</code></pre>
<figure class="post-figure">
<img src="https://www.mohitranka.com/images/blog/rdbms-vs-nosql.jpg" alt="Illustration comparing ordered relational tables with flexible document nodes" loading="lazy" width="1200" height="675">
<figcaption>
<span class="fig-caption">Datastore choice is a tradeoff about access patterns and operations—not a fashion contest.</span>
</figcaption>
</figure>
<p>I came across this quote in <a href="http://www.allthingsdistributed.com/2013/03/dynamodb-one-year-later.html">an article</a> while researching DynamoDB. I respect Werner a lot, but take his database-selection advice with a pinch of salt — he has a <a href="http://aws.amazon.com/dynamodb/">database</a> to sell.</p>
<!--more-->
<p>Most products never hit the kind of <em>scale</em> that relational databases cannot serve. RDBMSs have been around for decades and still work fine, including at <a href="https://www.facebook.com/MySQLatFacebook">large scale</a>. There are <a href="http://dev.mysql.com/">mature</a> <a href="http://www.postgresql.org/">open source</a> options, well-understood schemas, solid support for filtering and aggregation, big communities, plenty of people who already know the tools, default framework integrations, and a deep ecosystem of operational tooling.</p>
<h2>RDBMS are great for almost everything, unless…</h2>
<p>Relational systems are easy to work with. Monitoring, backup, and ops tooling is strong (and often free). They cover most use cases, help is easy to find, and — best of all — you probably already know them. Unless I have one of the requirements below, I would start with a single-node relational database.</p>
<ul>
<li><h3>Super high availability</h3></li>
</ul>
<p>The database server process is a single point of failure. If it crashes, or if you take it down for planned or unplanned maintenance, the system goes with it.</p>
<p>If you need something close to 100% availability, look at clustered relational setups or distributed databases that trade consistency for availability — Cassandra, for example. Until that requirement is real and written down, stick with the boring option.</p>
<ul>
<li><h3>Flexible schema</h3></li>
</ul>
<p>One of the great strengths of relational databases is the schema itself — modelling, <a href="https://en.wikipedia.org/wiki/Database_normalization">normalization</a>, and related ideas. You can design a solid model without knowing every query pattern up front, and it usually holds up. That said, <a href="http://en.wikipedia.org/wiki/Entity%E2%80%93attribute%E2%80%93value_model">not all data maps cleanly to tables and joins</a>. Postgres has some help via <a href="http://www.postgresql.org/docs/9.0/static/hstore.html">hstore</a> and <a href="http://www.postgresql.org/docs/current/static/functions-json.html">JSON</a>, but relational engines still prefer fixed schemas that play nicely with normalization. Flexible / EAV-style shapes <a href="http://karwin.blogspot.in/2009/05/eav-fail.html">tend</a> <a href="http://tonyandrews.blogspot.in/2004/10/otlt-and-eav-two-big-design-mistakes.html">to</a> <a href="https://www.simple-talk.com/opinion/opinion-pieces/bad-carma/">go badly</a> in an RDBMS.</p>
<p>If you genuinely need a flexible schema — and it is worth asking twice — look at NoSQL options. Do not pick them just because the schema feels slightly awkward on day one.</p>
<ul>
<li><h3>Horizontal scaling</h3></li>
</ul>
<p>If you expect data and traffic to grow beyond one machine, a single-node relational database will not be enough forever. There is only so much RAM and CPU you can throw at one box. Eventually you hit that wall, and then the choice is either a NoSQL system or a relational cluster.</p>
<h1>Epilogue</h1>
<p>NoSQL is not a panacea, and RDBMSs are still the right default for most applications. Unless you have a clear reason <em>not</em> to use a relational database, stick with one. Prefer the boring system until a real requirement forces something else — and write that requirement down so the next person does not re-open the debate from scratch.</p></content><category term="Blog"/><category term="data-platforms"/><category term="databases"/><category term="architecture"/></entry></feed>