Skip to main content

Command Palette

Search for a command to run...

From Blame Culture to Ownership Culture

Updated
•3 min read•View as Markdown
From Blame Culture to Ownership Culture
S

I have been working experience in areas of system administration, design, implementation & support of Windows Server Systems, Linux, Container and networking.

Production issue တစ်ခုဖြစ်လာတဲ့အခါ Engineering Team တွေကြားမှာ ကြားရလေ့ရှိတဲ့ စကားတွေရှိပါတယ်။

“ဒါက Development issue ပါ။”

“Code ကတော့ အဆင်ပြေပါတယ်။ Infrastructure issue ပါ။”

“QA မှာတုန်းက pass ဖြစ်ပါတယ်။”

“Requirement က ဒီလိုပေးထားတာပါ။”

“Security rule ကြောင့်ဖြစ်တာပါ။

ဒီအကြောင်းပြချက်တွေက technical perspective ကနေကြည့်ရင် မှန်ကောင်းမှန်နိုင်ပါတယ်။

ဒါပေမယ့် Incident ဖြစ်တိုင်း “ဘယ် Team ရဲ့အမှားလဲ?” ဆိုတာကိုပဲ ရှာနေကြပြီဆိုရင် Technical Problem တစ်ခုတည်း မဟုတ်တော့ပါဘူး။ Engineering Culture ရဲ့ Problem တစ်ခု ဖြစ်လာပါတယ်။

Production is a shared responsibility

Production Application တစ်ခုကို Development Team တစ်ခုတည်းက တည်ဆောက်ထားတာ မဟုတ်ပါဘူး။

Development က application ကို တည်ဆောက်တယ်။ QA က quality ကို စစ်ဆေးတယ်။ DevOps/SRE က deployment, infrastructure နဲ့ reliability ကို ထိန်းသိမ်းတယ်။ Security က system ကို ကာကွယ်တယ်။ Product/Management က requirements နဲ့ priorities ကို သတ်မှတ်တယ်။

ဒါကြောင့် Service တစ်ခုရဲ့ success ဟာ Team အားလုံးရဲ့ success ဖြစ်သလို failure ဖြစ်လာတဲ့အခါမှာလည်း Team တစ်ခုတည်းကို အပြစ်ပုံချလို့ မရပါဘူး။

ဒါကြောင့် Shared Responsibility ဆိုတာ Team တိုင်းမှာတာဝန်ရှိတယ်လို့ ဆိုလိုရင်းဖြစ်ပါတယ်။

Team တိုင်းမှာ ကိုယ့်အပိုင်းအတွက် clear ownership and accountability ရှိရပါမယ်။

Developer တစ်ယောက်အတွက် code merge လုပ်ပြီးတာနဲ့ ownership ဆိုတာမပြီးသွားပါဘူး။ Production မှာ issue ဖြစ်လာရင် troubleshooting နဲ့ resolution မှာ ပါဝင်ဖို့လိုပါတယ်။

QA အတွက် “Testing မှာ pass ဖြစ်တယ်” ဆိုတာနဲ့ မပြီးပါဘူး။ ဘာကြောင့် ဒီ scenario ကို မဖမ်းမိခဲ့တာလဲ၊ test coverage ဘယ်နေရာမှာ တိုးတက်ဖို့လိုသလဲ ပြန်ကြည့်ဖို့လိုပါတယ်။

DevOps/SRE အတွက်လည်း “Application issue ပါ” ဆိုပြီး ticket လွှဲပေးလိုက်တာနဲ့ မပြီးပါဘူး။ Infrastructure, logs, metrics, deployment history နဲ့ resource utilization တွေကနေ evidence ရှာပြီး resolution ရတဲ့အထိ ပူးပေါင်းဖို့လိုပါတယ်။

Security အတွက်လည်း “Policy အရ block ထားတာပါ” ဆိုတာနဲ့ မပြီးပါဘူး။ Security requirement ကို မလျှော့ဘဲ Business နဲ့ Application ဆက်လက်အလုပ်လုပ်နိုင်မယ့် solution ကို အတူရှာဖို့လိုပါတယ်။

Blameless doesn't mean accountability-free

ဒီနှစ်ခုကို မရောထွေးဖို့ အရေးကြီးပါတယ်။

Blameless Culture ≠ No Accountability

Blameless ဆိုတာ အမှားလုပ်သူကို ရှာပြီး အပြစ်ပေးဖို့မဟုတ်တာပါ။

Accountability ဆိုတာတော့ ကိုယ့်အပိုင်းမှာ ဖြစ်ခဲ့တာကို လက်ခံပြီး ပြင်ဆင်ဖို့၊ သင်ယူဖို့နဲ့ ထပ်မဖြစ်အောင် တိုးတက်အောင်လုပ်ဖို့ ဖြစ်ပါတယ်။

Incident တစ်ခုမှာ—

“Who caused this?” ဆိုတာ ထက် “What failed, how do we recover, and how do we prevent it from happening again?” ဆိုတာကို အရင်မေးသင့်ပါတယ်။

Leadership plays an important role

Blame Culture ဖြစ်လာခြင်းဟာ Engineers တွေရဲ့ ပြဿနာတစ်ခုတည်း မဟုတ်ပါဘူး။

လူတစ်ယောက် အမှားလုပ်တိုင်း အပြစ်ပေးခံရမယ့် Environment ဖြစ်နေမယ်ဆိုရင် Engineers တွေက naturally ကိုယ့်ကိုယ်ကို ကာကွယ်ဖို့ ကြိုးစားလာကြပါတယ်။

အဲဒီအခါ အမှားကို ဝန်ခံတာထက် အခြား Team ကို လွှဲချတာ၊ problem ကို ဖုံးကွယ်တာ၊ risky decisions မယူရဲတာတွေ ဖြစ်လာနိုင်ပါတယ်။

ဒါကြောင့် Leadership ရဲ့ တာဝန်က Accountability ကို တောင်းဆိုရင်း Psychological Safety ကိုပါ တည်ဆောက်ပေးဖို့ ဖြစ်ပါတယ်။

“Who made the mistake?” လို့ မေးတာထက်—

What happened?

ဖြစ်စဉ်တစ်ခုလုံးကို အချက်အလက်တွေအပေါ် အခြေခံပြီး နားလည်အောင် အရင်လေ့လာသင့်ပါတယ်။

What allowed it to happen?

လူတစ်ယောက်ရဲ့ အမှားကိုပဲ ရှာဖွေမယ့်အစား System, Process နဲ့ Control တွေထဲမှာ ဘာတွေလိုအပ်နေခဲ့သလဲဆိုတာ ရှာဖွေသင့်ပါတယ်။

Why didn't we detect it earlier?

Monitoring, Alerting, Testing နဲ့ Observability တွေမှာ ဘာတွေတိုးတက်ဖို့လိုသလဲ ပြန်လည်သုံးသပ်သင့်ပါတယ်။

How can we recover faster next time?

Recovery Process, Rollback Plan, Automation နဲ့ Incident Response Process တွေကို ပိုကောင်းအောင် ပြင်ဆင်သင့်ပါတယ်။

What should we change in our system or process?

Incident ကနေ ရရှိခဲ့တဲ့ သင်ခန်းစာတွေကို လက်တွေ့ကျတဲ့ Improvement Actions တွေအဖြစ် ပြောင်းလဲသင့်ပါတယ်။

Who owns the improvement action?

Action Item တစ်ခုချင်းစီအတွက် ရှင်းလင်းတဲ့ Owner နဲ့ သတ်မှတ်ထားတဲ့ အချိန်ကာလ ရှိသင့်ပါတယ်။

ဒီလိုမေးခွန်းတွေ မေးခြင်းရဲ့ ရည်ရွယ်ချက်က အပြစ်ရှိသူကို ရှာဖွေဖို့မဟုတ်ဘဲ ပြဿနာဖြစ်စေခဲ့တဲ့ အကြောင်းရင်းကို နားလည်ပြီး System နဲ့ Process ကို ပိုမိုကောင်းမွန်အောင် ပြုပြင်ဖို့ ဖြစ်ပါတယ်။

Strong teams don't avoid failure. They learn from it.

High-performing Engineering Team ဆိုတာ Incident မဖြစ်တဲ့ Team မဟုတ်ပါဘူး။

Incident ဖြစ်လာတဲ့အခါ တစ်ယောက်နဲ့တစ်ယောက် အပြစ်မလွှဲချဘဲ ကိုယ့်အပိုင်းကို ကိုယ်တာဝန်ယူပြီး အတူတူဖြေရှင်းနိုင်တဲ့ Team ဖြစ်ရပါမယ်။

ပြီးတော့ Incident ပြီးသွားတဲ့အခါ—

“Fixed.” ဆိုတဲ့နေရာမှာမရပ်ဘဲ — “What did we learn, and what are we changing?” ဆိုတာကို ဆက်မေးနိုင်တဲ့ Team ဖြစ်ရပါမယ်။

နောက်ဆုံးမှာ ကျွန်တော်တို့ တည်ဆောက်ချင်တာက Blame Culture မဟုတ်ဘဲ Ownership Culture ဖြစ်သင့်ပါတယ်။

Own your part. Support the team. Fix the system. Learn from the failure.

ကောင်းမွန်တဲ့ Engineering Culture ဆိုတာ ဘယ်သူ့ကို အပြစ်တင်ရမလဲဆိုတာ ရှာဖွေနေတာ မဟုတ်ပါဘူး။ Team အားလုံး အတူတကွ တာဝန်ယူပြီး ပူးပေါင်းလုပ်ဆောင်နိုင်တဲ့ Culture တစ်ခုကို တည်ဆောက်ခြင်းပဲ ဖြစ်ပါတယ်။

#EngineeringCulture #Accountability #Ownership #Teamwork

22 views

More from this blog

Application Deployment on AWS with Infrastructure as Code: Terragrunt, Atlantis & AWX - Part 1

Infrastructure-as-Code (IaC) ကို စတင်အသုံးပြုတဲ့အခါ Terraform က လူသုံးများတဲ့ ရွေးချယ်မှုတစ်ခုဖြစ်ပါတယ်။ Project နဲ့ Infrastructure အရွယ်အစား မကြီးသေးတဲ့အချိန်မှာ Terraform တစ်ခုတည်းနဲ့ Infrastructure

Aug 17, 20264 min read45
Application Deployment on AWS with Infrastructure as Code: Terragrunt, Atlantis & AWX - Part 1
V

Vital Tech Blog

27 posts