Posts

Showing posts with the label SRE

The Special Case of PID 1

 I was testing some things on our kubernetes cluster and wanted to quickly restart the k8s container without restarting the pod and having to wait for the graceful termination period. So I exec'd into the container and sent a signal to PID 1 ( kill -15 1 ) and expected the process to terminate while kicking me out of the container. Much to my surprise, it did absolutely nothing. I thought, hey, maybe for some reason it's hanging, let's give it a few seconds... A few seconds passed and still nothing happened. Okay, perhaps something is causing it to not properly process the signals, so I sent a SIGKILL ( -9 ) as those cannot be intercepted or processed in any way . Even that didn't kill the process. I was flabbergasted. Did we find a bug in the kernel, or did we find a way to make an unkillable process?   I decided to put a strace on the process from the host machine to see if there's any insight. At first I saw it was receiving the...

Making my Own Agent, Part 2: The Agentic Loop

Image
Prelude   In the previous post  I talked about the basics of an agent and why I decided to create my own. While I encourage you to go back to read the previous post, here's a refresher: an agent allows an LLM to interact with the outside world. I'm doing this for two main reasons, the first being that there is no great open source AI troubleshooting agent today and the second being that I find it helpful for my day-to-day activities.   I talked a little about the technical parts in the earlier post but it was all generic and pertained to all agents, in this post I am going to talk about how my AI agent works, from how it receives a signal to how it investigates to how it escalates or presents its RCA. Chain of Thought Traditionally, LLMs worked that you asked a question and it answered immediately by just spitting out an answer. Then in January of 2022 a paper came out detailing chain of thought (CoT), where you prompt the LLM to reason...

Making my Own Agent, Part 1: The what and why

I've been building my own agent harness for a few months and figured it was time to write about it. I don't have all the answers, but I learned a lot from other people's process posts and wanted to add to that pile. This is part one.  The foundation: You've probably heard the word agentic before, maybe a bit too much since it's sold as a marketing term. Agentic workflows, agentic coding, agentic hotel booking, agentic breathing. With as much hype as it's getting you would expect it to cure cancer and maybe one day it will but as of right now it's not creating any cures for cancer. So what separates agents and a chat application like ChatGPT? Well if you ask someone that today, the line is getting thinner and thinner. But a chat application is exactly what it sounds like where you chat back and forth with the LLM. With an agent you give it hands and fingers stretching out into the world to gather information and take actions. This can be any...