How to recognize Great Work

I drafted this post in early 2023 when AI was just getting traction. Scan it quickly and lets see if the premise still holds.

Evaluating performance isn’t about tracking hours or micromanaging, it’s about looking at two simple things: What you deliver and How you work with the team.

Both matter equally. Delivering great results with a toxic attitude doesn’t work, and being super friendly without ever finishing tasks doesn’t work either.

Here is what defines great results:

1. Great Outcome (The “What”)

Great outcomes mean delivering work above expectations without causing headaches down the line.

  • Done on time: You hit agreed deadlines without major delays or needing a rewrite.
  • Tested and code-reviewed: It actually works, and a teammate has reviewed it.
  • Documented and communicated: You explain what was built, why, and how to use it so nobody has to guess.
  • Clear value: The end goal or user value is obvious to anyone looking at it.

2. Great Behavior (The “How”)

How you show up every day sets the culture for everyone around you.

  • Team-first mindset: You care about the team’s overall success, not just your own ticket.
  • Strong communication: You keep people updated early and often—especially crucial in a WFH setup.
  • Solution-oriented: You follow a simple loop: Plan -> Execute -> Refine.
  • Real ownership: You take responsibility for your work, and you’re genuinely glad to jump in and help others when they get stuck.
  • Learn and lift: You pick up new things fast and share that knowledge with the team.

What still holds today in the AI era?

What Has Shifted in the “What” (Deliverables)

Today producing outcomes is GenAI salt and butter. AIs own training and harness have huge impact on outcome. It’s really good at executing, drafting and prototyping, but it lacks human judgement and accountability. Lets see how outcome goals changed in AI era.

“Tested and code-reviewed: It actually works…”

  • The Reality: AI generates plausible, syntactically correct code that can contain subtle edge-case bugs, security vulnerabilities, or anti-patterns.
  • The Shift: Writing the code is no longer the main bottleneck; verifying, validating, and reviewing it is. High performance here isn’t just using AI to write unit tests—it’s having the domain knowledge to know what edge cases the AI missed and ensuring system architecture isn’t deteriorating under the hood.

“Clear value: The end goal or user value is obvious…”

  • The Reality: AI can generate code, draft documentation, and outline features based on prompts, but it doesn’t understand product vision or business context.
  • The Shift: The human role shifts from builder to curator/director. Defining the right problem, validating that the solution actually solves a real user pain point, and trimming AI-generated feature bloat are distinctly human responsibilities.

“Done on time…”

  • The Reality: The baseline speed for standard tasks has increased across the board.
  • The Shift: Hitting deadlines is less about manual keyboard time and more about task decomposition, prompt engineering, and rapid iteration loops.

The “How” (Behavior) is non-negotiable

Culture is human: AI cannot make for psychological safety, facilitate cross-functional discussions, mentor a junior engineer or align across teams when priorities clash.

Communication in remote or hybrid teams: Since AI speeds up code generation, async sharing and proactive updates matter more. A team coding 3x faster without tight communication leads to 3x chaos faster.

The “What” vs. “How” balance

High performance still requires both pillars. Generating 10 pull requests a day with AI while bulldozing team processes is just automated toxicity.

For the reference where this method is in use across top companies

Company Name ProjectApplication
NetflixPerformance vs. Values Matrix & “No Brilliant Jerks” PolicyEvaluates technical delivery alongside cultural adherence. High delivery combined with toxic behavior gets zero tolerance.
GoogleProject Oxygen & “What & How” Rating ScaleUses 360° reviews to rate tangible output (“What”) equally alongside collaborative behaviors, coaching, and communication (“How”).
StripeEngineering Operating Principles & Leveling GridRates engineers on Execution Rigor (code quality, testing standards, documentation) paired with Team Impact (ownership and mentoring).
GitLabAsync-First & Values-Based EvaluationEvaluation assesses performance through documented output, asynchronous communication rigor, and self-directed ownership in a distributed environment.
AmazonLeadership Principles + Delivery MetricsPerformance reviews split focus 50/50 between hitting concrete delivery targets and demonstrating behaviors like Ownership and Bias for Action.
General ElectricJack Welch 2×2 Performance GridPioneered the classic 2×2 matrix evaluating staff on two independent axes: Results (Outcome) vs. Values (Behavior).

Summary

High-quality deliverables combined with a strong collaborative culture define high-performing engineering organizations.

It holds in AI era too, but require a subtle shift in your works processes. It makes the work more demanding by extra skills like project management, QA assistant and architect, but it also makes the work more interesting and comprehensive.

The Chaos Monkey Controversy: Why Overuse is Killing Developer Productivity

In the world of Site Reliability Engineering (SRE), few tools have sparked as much debate as Netflix’s Chaos Monkey. What started as an innovative approach to building resilient systems has become a lightning rod for controversy, dividing the engineering community into passionate camps. After years of watching teams struggle with its implementation, I’m convinced that the tool’s widespread misuse has created more problems than it solved.

The Promise vs. Reality

Netflix introduced Chaos Monkey in 2014 with a simple premise: randomly terminate instances in production to force teams to build more resilient systems. The idea was sound – if your system can’t handle the unexpected failure of a single component, it’s not truly robust. But somewhere between Netflix’s careful, strategic implementation and the broader industry adoption, something went wrong.

The evidence is everywhere. Reddit discussions in r/sre are filled with horror stories of cascading failures triggered by poorly timed Chaos Monkey deployments. HackerNews threads document teams spending more time recovering from simulated failures than actually improving their systems. As one frustrated developer put it: “Chaos Monkey causes chaos, it does not fix it.”

The Case Against Overuse

The critics’ arguments are compelling and backed by real-world evidence:

Reduced Developer Productivity: Constant, unpredictable deployments disrupt development workflows. Teams report spending excessive time rolling back changes and restoring services instead of building new features. When your developers are in perpetual recovery mode, innovation suffers.

False Sense of Security: Perhaps more dangerously, frequent disruptions create complacency. Teams become so accustomed to recovering from simulated failures that they lose focus on preventative measures. The perception of resilience becomes skewed – surviving artificial chaos doesn’t guarantee handling real-world incidents.

Operational Overhead Explosion: Every Chaos Monkey deployment requires investigation, analysis, and often rollback procedures. This consumes significant resources that could be better spent on proactive improvements and strategic technical debt reduction.

Misinterpretation of Failure Data: Controlled, short-lived interruptions bear little resemblance to sustained, complex outages. The data collected during Chaos Monkey tests cannot reliably inform long-term architectural decisions because the test environment fundamentally differs from genuine failure scenarios.

The Netflix Model: What Actually Works

Netflix’s original implementation wasn’t the blanket deployment strategy many teams have adopted. Their approach was surgical, strategic, and tied to specific maturity milestones. They deployed Chaos Monkey on reduced scales, focused on services that were already stable, and closely monitored the results.

The key difference? Netflix treated Chaos Monkey as a precision instrument, not a blunt force tool. Their engineers understood that the goal wasn’t to cause disruption – it was to observe system behavior under stress and identify specific weaknesses.

The Middle Ground: Disciplined Chaos

This doesn’t mean Chaos Monkey is inherently flawed. When used correctly, it can provide valuable insights. The critical factors for success include:

Service Maturity Requirements: Deploy only on stable, well-understood services. Using Chaos Monkey on immature systems is like stress-testing a house of cards – you’ll learn it falls down, but not much else.

Controlled Frequency: Deployments should be carefully scheduled and spaced, not constant background noise. Teams need time to implement improvements between tests.

Clear Success Metrics: Define specific learning objectives before each deployment. What exactly are you trying to discover or validate?

Comprehensive Monitoring: Ensure you can observe and analyze the complete impact of each test, not just surface-level metrics.

My Take: Strategy Over Chaos

After witnessing both spectacular failures and genuine successes with Chaos Monkey, I believe the tool’s value lies in its strategic application, not its frequency of use. The teams that succeed with chaos engineering treat it like any other engineering discipline – with careful planning, clear objectives, and disciplined execution.

The controversy surrounding Chaos Monkey ultimately reflects a broader issue in our industry: the tendency to adopt powerful tools without fully understanding their appropriate context and limitations. Netflix built Chaos Monkey for their specific needs, scale, and organizational maturity. The problems arise when teams apply it blindly without considering their own unique circumstances.

Looking Forward

The evolution toward tools like Chaos Gorilla and more sophisticated chaos engineering frameworks shows the industry is learning. We’re moving away from random disruption toward targeted, hypothesis-driven resilience testing. This represents a maturation of the chaos engineering discipline – from “break things and see what happens” to “systematically validate our assumptions about system behavior.”

The lesson isn’t to abandon chaos engineering, but to approach it with the same rigor we apply to any other engineering practice. Used strategically, tools like Chaos Monkey can strengthen systems. Used carelessly, they strengthen nothing but frustration levels.

The choice is ours: embrace disciplined chaos or remain victims of chaotic discipline.

Tech Team Leaders’ Guide to Strategy

Building a tech strategy is a core responsibility of the CTO, VP of Engineering, or Head of Engineering. Involving team leaders in this process ensures a more grounded and effective approach.

Tech team leaders play a crucial role by defining roadmaps for their teams, which, in turn, provide the foundation for an effective high-level strategy. To achieve the best results, continuous collaboration between leadership and team leaders is essential.

Let’s explore how to create a roadmap that is both practical and aligned with the company’s overall vision.

Building an Effective Team Roadmap

A team roadmap is a strategic document that outlines product needs, infrastructure requirements, modernization efforts, and compliance and security considerations, among other critical aspects.

An effective roadmap goes beyond listing high-level initiatives or goals. It expands on each goal using the Diagnosis, Policy, and Actions framework, helping to answer the Why, What, and How of every initiative. This approach fosters trust, alignment, and transparency with top-level leadership.

The Diagnosis, Policy, and Actions framework, developed by Richard Rumelt, consists of:

  • Diagnosis – Defining the problem that needs to be addressed
  • Policy – Establishing guiding principles and constraints for the solution
  • Actions – Defining concrete steps to implement the solution within the given policy

Let’s explore a few examples.

Example 1: Modernizing the Infrastructure

Diagnosis: Our current infrastructure relies on outdated and proprietary components, leading to scalability challenges, high maintenance costs, and slow adoption of new technologies.

Policy: Prioritize open-source and cloud-native solutions for new developments. Maintain legacy systems where necessary but avoid further expansion of proprietary technologies.

Actions:

  1. Identify and replace critical proprietary components with open-source or cloud-native alternatives.
  2. Standardize infrastructure automation and provisioning to improve scalability and maintainability.
  3. Update internal documentation and on-boarding materials to reflect new infrastructure standards.

Example 2: Upgrade the Database

Diagnosis: The current database version has reached end-of-life and is no longer receiving security updates or feature enhancements. An upgrade is necessary to maintain security, stability, and performance.

Policy: The database upgrade must be performed with zero downtime to avoid service disruptions.

Actions:

  1. Test new database version in the QA environment to ensure compatibility
  2. Create a full backup of the existing database.
  3. Implement a Blue-Green deployment strategy to minimize risk during the upgrade.
  4. Communicate the upgrade plan and schedule a rollout window.

Example 3: Improve Cloud Cost Efficiency

Diagnosis: Cloud expenses represent a significant portion of overall costs. Unused or underutilized resources contribute to unnecessary costs.

Policy: Optimize cloud usage by right-sizing instances, using auto-scaling, and enforcing cost-control policies.

Actions:

  1. Conduct an audit of cloud resources to identify inefficiencies.
  2. Implement auto-scaling policies for workloads with variable demand.
  3. Use reserved or spot instances for predictable workloads.
  4. Set up monitoring and alerts for unexpected cost spikes.

Conclusion

  1. Structuring your team’s roadmap using the Diagnosis, Policy, and Actions framework ensures clear prioritization and alignment with the company’s overall strategy.
  2. This approach facilitates productive discussions with top-level leadership, leading to better decision-making.
  3. It improves transparency, trust and accountability across all levels of the organization.

Have you faced challenges when implementing a strategic roadmap? How did you overcome them? Drop a comment below and let’s learn from each other!