Video: Your Homegrown Feature Flags Worked, Until They Didn't | Duration: 1792s | Summary: Your Homegrown Feature Flags Worked, Until They Didn't | Chapters: Welcome and Introduction (1.115s), Host Introductions (82.335s), Scaling Flag Challenges (195.83s), Rachel's Case Study (330.115s), Feature Flag Failures (458.66s), Live Feature Demonstrations (653.795s), Progressive Rollout Strategy (980.275s), Emergency Rollback (1188.375s), MCP Integration Demo (1374.475s), Flag Management Scale (1634.435s), Next Steps & Closing (1690.075s), Closing & Thank You (1761.62s)
Transcript for "Your Homegrown Feature Flags Worked, Until They Didn't":
is. Oh. Alright. I think we're on. We're gonna start welcoming folks to the room. Good morning, good afternoon, good evening, everyone. When you get a second, let us know where you're logging in from. We always have such such a variety on these events. Alright. Alright. As we give, a minute or two for for folks to join, let me go ahead and introduce the topic we will be covering today. Thank you so much for joining us, for this feature flags webinar where we discuss homegrown feature flag system that are very prevalent, among pretty much most org engineering orgs, anywhere between 40 plus, 50 plus, 60 plus percent is what we've been seeing as we've been running different, surveys and research and researches. Let's see how many we have and where we are. Alright. With that, we can start, moving on to introducing our host and drivers for today, Avan from product and Sean from the sales team. Let's go ahead and do that. So hi, everyone. My name is Shelley Eisen Liebne, director of product marketing in building products, taking them to market for over twenty years now, and over the last decade, really specialized in technical audiences and technical value. I'll go ahead and conquer it over to Ivan. Yeah. Thanks, Shelley. Hi, all. My name is Evan Malter. I'm senior director of product management here, driving our feature management release areas in Cloudbees. And I've been, working in the software delivery, release, management, feature, plugging space for quite a while now, ten plus years. And really happy to be here to talk to you all. I'll pass it to Sean. Thanks, Ivan. So I'm an intelligence solutions architect. Before that, I was a senior SRE at Cloudbeads. And for ten years before that, I was in DevOps with IBM, which was really just more SRE. It was it was focused on the SRE aspects of DevOps. So a lot of the things we'll be talking about today are things that I lived for the last decade and a half at this point. Alright. Back to you, Shelley. Alright. Awesome. So with that, a few housekeeping real rules before we jump right in. If you have any questions, feel free to drop them in the chat or the q and a section. We will try and answer whatever we can, in the chat or at the end, if we have some time for q and a. If we are not able to cover all of these, then we will follow-up with you right after. And without further ado, I will hand it over to Ivan. -Great. Thanks, Shelley. Yes, I'm really excited to get to talk to everyone here today. So we are talking about a problem that we've heard often when we're speaking with our customers, with folks in the industry who are working to scale their, feature flagging practices. They may be small businesses all the way to large and enterprise, and we see a common challenge repeated. They want to start bringing more control to their releases through feature flags. And when they're getting started, to them, it makes sense to build a feature flag system themselves. They don't wanna get locked into a vendor. It's easy to build and ship in, you know, a matter of weeks and actually even faster now with AI. Right? It's a homegrown flag solution can be spun up in just an afternoon. So it's it's easy to get started. Build it yourself. Start introducing configurations and flagging into your software. But it's really, after this starts to grow with time, that they start facing challenges. So the first one here is when, when AI code volume starting their system. Right? The growing number of feature flags require more governance, more control. And with AI code generation, it can double the rate of creation of not only code, but the number of flags that are being added to your system, creating a larger challenge with the proliferation of flags throughout your code, as well as managing, cleaning up still flags, all of that becomes amplified. Second, there isn't really discipline or policy that enforces best practices with that scaling of flags. And so without that hygiene, the management, the control becomes even harder. And then last is the lack of governance and traceability. You you know, how do you find out when an issue occurs? Who made that change? Who where did that flag configuration change come from? But having that full traceability and understanding, and also having controls such as our background who can make a change so that you don't have errors or changes that shouldn't be made to flags coming into your system, especially in your production environment. So moving on, we wanna talk through a real life scenario, that that's similar to what we've heard repeatedly from our own customers as we're talking to folks in the industry. So we're gonna follow the story of Rachel. You may relate to her. She's a director of platform engineering. She works for a company with 600 people, 240 engineers. And they have 470 active flags in their system. It is a homegrown flag system that's been active for three years. So the this number 475 is a lot. It's been growing and it's been growing even faster with developers now using agentic coding as part of their processes. So this may sound familiar to you all. It's not really a unique case, right? In our own surveys, we've asked and we see that 48% of teams, are on homegrown flagging systems, 47% say they are using flags and more release controls, across their release processes. So Rachel here, she wasn't you know, she started with a small problem. She built a homegrown, their own solution for future flagging, and thought that was an easy path to get this problem solved. But now she's realizing that it wasn't in her plan to actually be managing a feature flagging platform. It grew organically over time and is now hitting the limits and starting to feel some pain. So let's see where this starts breaking. The strain really starts to manifest in different ways. We we see it in these solutions often where they don't scale. They become less reliable and introduce more risk. Let's make it a little more concrete and look at what some scenarios that Rachel's team is facing. So they're shipping a new pricing and checkout experience in their application, and it's in one release where they have 12 flags, many changes. This is alongside the 400 plus flags they have in their system, 200 plus engineers who are touching the code base. So let's see what happens after they've released this. So first here, it's Friday evening, 04:47PM. Release is done, but now they're seeing errors in the checkout system. They're spiking. And so now it's the hunt where which flag caused the issue? Why why are we seeing errors? It takes tens of minutes to find out what happened, and it to find the trace back to what change was made, what which flag is it tied to. And what it turns out is a flag that was actually not even supposed to be enabled in production. It was meant to be tested in the staging environment, was accidentally enabled in prod. The flag system that they have doesn't support RBAC. Anyone who can reach it can change a flag in any environment. No guardrails exist, and there's no authorized approval that's required for production configuration. So somebody can make this mistake without even realizing it, and they only start seeing it when these errors occur in prod. So they found the issue. Let's turn it off. Right? But now it's the next problem. The admin UI is down. No one knows the no one knows, you know, who's available right now can make that fix because the the team that maintains this is not really available, so they decide to redeploy all the changes. And this takes an additional twenty two minutes to redeploy, the changes, roll back. They lose the other changes that were part of this release as part of that process. So now it's a week later, and 3AM. So they an agent is working on that same pricing service. It reads the code, sees that there's a conditional block, makes an experimental change to pricing without seeing where what that flag configuration is from from the agent itself. So they don't see it doesn't see that the flag is live for 40% of the production traffic, and now they're seeing the wrong price in production. So nobody caught the change. It ended up in front of users, only somehow caught at 3AM. And now we're in the same process of investigation. Let's find out what happened, trace it back, and, undo the change. And now we're at the the last one here where it take where at the end of the quarter, auditors are asking what happened when that error occurred. How you know, let's find out exactly who made the change, why it was made, and that process itself now going back in time, trying to find those details, it takes a day plus to find who was the player who introduced the configuration change. Was it a human? Was it AI? It's not even available in the data they have available. I'm going to hand it off to Sean in a moment who's going to talk about how these can be addressed with Cloud based feature management. We have our back and, kill switches that can roll back changes and issues in less than a minute, in a consistent, reliable fashion. Agents have access and can read live flag states. So, changes are informed and can be done from anywhere to the flag configurations. And every change is attributed, whether it's a human or agent. There's automatic, auditing happening over time collected that can be read back at any time. So, Sean, let's see this in practice. Thanks, Yvonne. Let me start sharing here. So we'll talk about Rachel's story. She's the director of engineering, but she also kind of serves as the release manager. So she's the only one that can actually change flags. Let's see what that looks like. So here, over at the left, we have all the default colors of the header theme. So let's see if we can look at these and change the default to dark mode there. So here, with real time evaluation, you see a change in less than five seconds. You see the flags all turn into that dark mode color. There's no redeploy, no code changes. We didn't even need to refresh the page. So that's what Rachel can do. But now let's look at one of her teammates. So here, we see her teammate, Chris, who maybe is not as trusted with changing flags in the environment and needs to go through the approval of our lease manager, Rachel. So when he looks at the same thing, what he's going to see is this UI where it just shows request approval up in the top. So maybe he thinks that enterprise should get a nice vibrant color. And then if it's not enterprise, maybe if they're in The US East Region, let's make it dark. And then if it's neither of those, let's just make it default is what Chris wants to do. So he requests approval. We're gonna see that there, and it'll get routed to Rachel. How does that RBAC get managed in our system? So you can actually create the roles yourself. You see here, we have that typical sort of RBAC matrix where you can assign these to user groups, to specific users, and handle all of those. What does that look like over for Rachel? So let's look at this, and we can actually see in her UI exactly where those are awaiting approval. So she has a queue of flag changes that people have requested that they want to make, and she wants to try to approve these if they look good and then verify that she's happy with what the UI looks like. So she has something in the back end here where she can actually simulate traffic and will see failures in real time. So we'll turn this on. It's sending traffic, and it'll track the last failure that that or I should say the first failure since she has reset. So for the sake of this session, let's go ahead and reset. We can keep an eye on this dashboard, and let's start with the change that we just got from Chris. So header theme. So we look at this, and we can go through. We see for enterprise, it's Vibrand. For US East, it's dark. Otherwise, it's it's default. So over here, if we look at each of these, this is meant to represent an enterprise customer, so 1,200 employees. This is meant to represent mid market. They're in Massachusetts. And then finally, we have Tidepool in California, and that's the small company. So here, we'll say, looks good, and approve that. And, again, with real time evaluation, we'd expect to see each of these change based on the flag evaluation. And over here, we see that enterprise does look vibrant. We see that the East US East customer shows dark, and then we have the default for falling through that logic. So it's the same code, same deployment, but everyone gets their own unique experience. But it's a little cliche, right, showing just header colors changing. So let's see if we can look at some actual meaningful features here. So for this recall adviser, this is intended to be an AI powered chatbot. So maybe we just want to request this for the power users. So the change requested here, it says if the target group matches enterprise, then show them the chatbot. Otherwise, hide it. We don't wanna show it otherwise. So, yep, that's what we want. Sounds good to me. So let's see what that looks like. So we can approve that. This is supposed to be our simulated enterprise customer. They have a new chat button here. If we look at the others, there's nothing down there. Nothing down there. So the only one that has access is this customer. And we have a little message here. Let's say what is elastorb? Oh, the improper key wasn't configured. So maybe this is an issue where it's something the developer is still working on. Maybe it's just an issue where the, the key should have been updated and wasn't. But either way, Rachel sees that. And right here, she can just turn it off. And then back here, the chat button is already gone. So we don't see it. We don't have to have to refresh anything. Same deal as as the header color changing with a real feature here. As the next one, let's look at calendar view. So here is one. One someone on Rachel's team requested this. What they said is that it's actually a very re resource intensive flag. So they're worried about what the back end is going to be able to do for enterprise customers. So for small customers, they decided enable it for all of the small customers. For our mid market, give half of them the feature, but not the other half. And then for enterprise, they actually have a scheduled deployment here. So you can see over the course of five days, it will continuously scale up from zero to a 100 to zero to zero zero. I'll tweak that so that we can actually well, that's fine. So here we could would see zero, five, ten, twenty five, 50, and a 100%. And what that would practically look like here, if we go to the progressive rollout so as we scale up which users are going to get it, it uses an algorithm against those users and will actually just add users. So it's only additive. You can see the simulated volume here of what that would look like. And then over over the days, we went up to 10%. You can see the the volume down at the bottom. It looks a little larger. It's increasing there. 25. We see more users getting selected. You'll note that no user ever loses the blue box once they once they get it. So you're not going to be scaling for customers and then they lose the feature as you as you start to scale up and make it available to more customers. And then finally, once the feature is completely ready, we see that it's all on the new feature here. It's a 100%, and that fills out. But let's go ahead and approve this and see what it looks like. Okay. So it should work for a 100% of our small users, and we now see calendar here. So we can click it. It says calendar view is disabled. This is not intended, but worked out. So this is to show that it's actually something that can be blocked on the front end or the back end. So this feature, we saw it pop up in the web UI. It was there. Let's go ahead and just try clicking it again. Now after I had another second or two to, reconnect in the back end and everything, we can connect it. So it's not just a a feature that someone can go in and and change the UI to bypass the feature flags. It's something that we can actually enforce them back into, which is a good example of that. Here, this is our mid market. There's a 5050% chance they have this. So based on based on this, they must be in the second half. So let's try bumping this up to 75% and see see what that looks like for them. So if we say 75% chance, can we see it update for them? Still not yet. There they are. They're between 7590%. So there you have it. They can also see the calendar. And then enterprise hasn't scaled up at all yet, so it's not available. So with that, we can actually perform we can actually monitor the performance and slowly scale it up to check back end performance. Let's move on to the next change, which is dashboard redesign. So this one, you see it has a lot of code references. It sounds, sounds like it's probably a big change based on the fact that it is dashboard redesign. But if we look in it, it looks like there's there's a lot going on here. We have a lot of conditions. Let's say that Rachel reviews these. She approves. She says, yeah. It'll be fine. It looks it looks good. Let's let's say that's good. Now here, do we see a new dashboard? So it still looks the same. Maybe it's something that's done on login. So if we try that, we get an error message. So let's go back to that, health dashboard that we created that, Rachel was monitoring. You can actually see each of these black lines is the only change flags. And since this last change, it looks like it's all errors. So based on what we've seen, we can probably surmise what that is. Right? We we think that it's the it's the dashboard redesign, but let's take this and see if we can actually follow this through in our audit trail. So twelve eighteen, I wanna go back here and look at our audit history. So looking here how did I say? 18. Interesting. The calendar view caused an error too. Okay. So from here, I wanted to look at the the dashboard redesign, and we can see that it had all these conditions added. And we can surmise that the the errors are caused by that. So let's go back to the actual feature management there. And when you're looking at the flag, from here, we see it has all of those conditions. There's some regex. We don't want to accidentally, like, lose all of those. So from within the the environment, we can disable it right here. There's just this emergency kill switch that will immediately turn it off, fix the issue. It's also available from the top level for each environment. So we can just say, regardless of what the conditions inside it are, I don't want you to surface flag. It should always just be false. So based on that, let's look at our page again. We see the flag change was made right here, and we actually see that requests are starting to look good again. So we're getting blue. Can we actually sign in now? Okay. Great. So it looks like the issue was addressed. We saved the problem at 04:47 in record time. It took less than thirty seconds to check it and then roll it back. Okay. One last thing. I want to share that everything I did in the UI is also available on our MCP server. So in cloud, I can just ask, can you tell me about my feature flags in the dev environment? So we'll give it a second to do this. While we're letting Claude think, I want to take a minute just to show how how simple making these changes can be. So here is a change from start to finish that we could potentially have. The middle block here, the middle def is the actual change that you see. In here is the if logic where you basically just say, hey. If the flag is is true, I want you to show it. Otherwise, don't. And then all you need to do to register it with the SDK is initialize the variable in one file. So as soon as you build this, deploy it, your flag will appear in the Cloud Bee's UI, and you can start manipulating it immediately. So rather than wait for this to come back, I asked the same question before this meeting just to make sure make sure it, would be ready if it was taking too long. So here's the what we saw come back is the table showing the definition of each flag, showing the actual description of what the rules are. You'll notice for calendar view, it actually shows that the default rollout ramps from zero to a 100% from eight twenty six to nine two where it actually intelligently described it based on, based on the schedule ramping up. So with that, I'm going to let Shelley take over the sharing again, and we can revisit the end of Rachel's day. Yay. Thank you so much, Sean. That was really helpful with all the illustrations of what to actually to expect in terms of what's happening for users, for the business, the risk, the way Rachel and Tim can manage it, but also how this impacts the developers, on the other side. So, let us let me share my screen again. And, Yvonne, I will ask you to bring us home. I let's get out of mute quickly. Sorry about that. Yeah. Thanks so much, Sean. I think it was great to, see it in action, feel a little bit of what the the pain that we talked about, how easy it can be when you have the right system and platform in platform in place, with all of the governance, and audit capabilities to help you along the way. Right? So we saw what Rachel's Friday was before at 04:47PM. Not a lot of fun. But with this kind of solution, it's just one toggle. Right? In under sixty seconds, you can recover from an issue that comes in production. The audit log is captured automatically along the way. So it's not a hunt to find what flag, what it in introduced the issue, you can easily map it to the time of your incident and find what was the flag change, who made the change, who was the actor behind it. We saw in the audit that you could attribute it to the user or even the, AI agent that is introducing that change. There's no one, you know, cloud users here, maintaining the reliability of the system. You don't have anyone on call rotation for, the flag configuration. I'm so sorry. I my bad. It wouldn't go back. Nope. One second, guys. We do have just a couple of minutes, so let's make sure, we end on time for the folks. Apologies for this. Alright. Let's go again. Alright. Let me know when you're ready. Yeah. No. So the and as I said, no one, you know, no no one is focused on maintaining the flag system. Instead, they're focused on delivering their software and the quality of that software. They are able to set up the blast, the roll progressive rollouts to control blast radius. So Rachel and her team are not maintaining the FutureFly platform anymore. Instead, they're just getting to use one and bring the value to their customers in the right way. So moving on, I won't go through each of these on the next slide in detail, but there are six things that we find, you know, need to be true for flag management to work at scale. We want every flag to have an owner. We want progressive delivery. AI agents should get to be able to see what's happening, understand, and make changes directly. And you have hygiene and auditability. So with that, Yeah. I think we're ready to close. I'll pass it back. I don't know if we have time for questions or we Yeah. We actually have, like, maybe a minute. So if anything comes in the chat, we will address it shortly. But I will take us to our final slide, which is what are the next steps? Where to go next? So, Ivan, feel free to take it away. And if any questions come in, we'll address them. And if not, we're happy to answer any questions in another time. For sure. Yeah. Please, you know, do contact us if you have any thoughts or questions, but what we would love to see you, see you guys come and check out our feature management workshop shop where we can get you started. And you saw how Sean shared how easy it is to, just introduce a flag. It's very quickly to integrate the SDK in one step, create a flag, really actually flip that kill switch, see it in the environment, try out the targeting rules, and say goodbye to a full redeploy. Do this in a quick session with us where you can really feel and understand the value of heat feature management from CloudBees, with a lot of ease. So please contact us, and we'd love to work with you towards that. Alright. Thank you so much, Sean. Thank you so much, Yvonne. Thank you, everyone, for joining us today. You will get an email with a recorded a recording and follow-up assets, and we would love to engage with you, see where you are in the in your journey of feature flags, what are your biggest pain, and and how we can help. Have a wonderful rest of the day. Bye, everyone. Thanks, Sean. Alright. Thank you.