Ten years ago I was starting the co-authoring our first book, which turned out to be the first published on Oracle PaaS. Specifically it was about Integration Cloud Service, which was became Oracle Integration Cloud (OIC). The book was published in the Spring of 2017. It was a rewarding experience, and the first of a series of books written over the last 10 years.
Most technology books, particularly those that are focussed on a specific technology tend to have a relatively short shelf life. – part of the reason publishers have early release programmes.
But our first book has been promoted to us as part of Amazon’s daily Kindle Reads email as you can see here…
While I am certainly not complaining, I am curious as to how Amazon’s algorithm choice the book.
As the OpAMP project has grown from a small idea to something pretty substantive, we’ve decided to put it in its own Org and repository with GitHub. This should mean it no longer looks a pet project. The repo name has also been changed as we have address more than just Fluentd and Fluent Bit needs with the Elastic observability pipeline tools having been incorporated with the basic features as well asl Vector.
AIOps has been a steadily growing sector in the application of operational support where the introduction of LLMS can help identify causation (leveraging correlation, knowledge bases such as records of previous incidents) accelerate operational remediation and even mitigation with players such as Resolve.ai, ObserverIQ, and Keep (Open Source).
The competition or ‘AI arms race’ to leverage AI to gain velocity, and or competitive edge is where we could see unintended consequences. A little bit of back ground first.
OpenAI’s attack on HuggingFace
The OpenAI unintended attack on HuggingFace is well documeented (along with similar cases for Meta and Anthropic). Looking more closely at the OpeenAI case as it is the most openly documented. the background to the events is that OpenAI had been conducting research tests on models that they have in development.
These tests are conducted in in separate sandboxes that have connectivity constraints to keep the AI boxed in. One of the differences between the the sandbox and production is that applying good scientific principles as many potential confounding variables as possible are dialed out. Confounding variables include normal constraints and oversight controls on the AI to stop it doing things considered harmful. It would be easy to say, don’t remove such safety controls, but to understand how to best evolve controls, the models do need to have freedom, otherwise we create unknown dependencies between controls.
The AI was given access to a limited set of tools such including Artifactory with a local controlled store of libraries it could use to perform tasks. Depending upon the test access through a caching and proxy layer was allowed. Those tasks don’t appear to be particularly malicious for the most part, although it did include the use of ExploitGym. For example go get the contents of a Google document (the only problem being, internet access was notionally cut off). Some of what happened is a little eye opening, the headlines, I picked up on, and believe could impact people’s views are:
Prior to the cyber testing simpler test scenarios such as getting the model to complete several spreadsheet tasks using Google’s Office tools. Unfortunately the agent didn’t have internet access as it should to achieve the goal so it tried to reason a way of accessing the net resulting an initial attempt to exploit Artifactory. While it failed, it did store content into Artifactory.
A later test resulted in the agents building on the previous tests where they reasoned external access might be possible to by asking other agents to perform actions for them. This lead to Artifactory being used as a collaboration channel.
Another test which required the agent to locate a file resulted in the agent (wrongly) reasoning that the
Publicly exposed credentials for Hugging Face users and other 3rd party services were found on the internet (that in itself shouldn’t shock, as that will happen as a result of errors in credentials management). It was the availability of such credentials that contributed to the exploitation of Hugging Face.
The AI reasoned that to access the internet it needed to a way out. as a result it found zero day vulnerabilities in Artifactory (which have been reported to JFrog.
The following timeline – tells the story over time.
How does this impact AIOps?
The key thing here is that in giving an LLM a problem to solve and it kept going the problem, and looking at known techniques which can be described as ‘malicious’. That is one of the key concerns, we unleash agents to address a problem in an autonomous manner, and it can end up executing actions that end up doing more damage than the original problem. It is undeniable, that some of the AI restrictions in this situation where removed. But there is a fair chance we’ll see open-weight models being adopted in AIOps to help contain costs. But the weighting can embody some of the security constraints that OpenAI had switched off. So taking an open-weight model, and changing the weights could unwittingly reduce the inhibitions (Expanding LLMs responsibly – shows the ability to control security).
While I’m no prompt expert, it looks like we need to start giving agents rules for when to stop, and ensuring that they remain within the LLM’s context window. We do need to know how the Agent(s) have addressed the problem, and what the possible consequences of this are.
It is human nature to trust things if the out come looks correct. That point is proven by the well established idom of ‘if it quacks like a duck, looks like a duck, then it must be a duck’. This means we’re at risk of trusting the AI has got the solution correct. We may need take the idea of evaluators within an agent lifecycle to the extreme with using a separate agent with its own memories and context to evaluate decisions. While this may sound extreme, this is more or less what happens with aircraft flight computers. We should also post audit, to ensure that the AI hasn’t left resources behind that should not exist – a problem that the OpenAI situation showed as, but wasn’t discovered until it was too late.
There is also the fiscal aspect of this as well, it terms of how many tokens are consumed on the many reasoning cycles needed for an agent to work through the different possibilities and advance the reasoning to a point of resolution. Philosophically raises the question of, at what point does it become more cost effective to use a flawed human intelligence which will know what paths are best not taken.
Ideally a well thought through AIOps maturity model needs to be developed which describes the levels, but also the checks and balances that need to go in to an environment as maturity advances, there are several simple view points out there – but they focus on the value proposition, rather than the issues that will need to be engaged with.
It is funny how when websites reference S3 (Simple Storage Service) storage, they typically reference a couple of the Hyperscalers (usually from Azure, GCP and AWS). While I get the reference to AWS, as they developed the service that has become the foundation for the API specification. But the association with just those hyperscalers is deeply misleading (and, to be honest, it bugs me). Any infrastructure as a service (IaaS) provider provides an S3 solution, and would probably not be taken seriously if they didn’t.
But S3 is not just a cloud-only capability today; enterprise-class SANs have S3 as part of their offering. There are open-source offerings that will allow you to overlay S3 onto commodity hardware Ceph, MinIO, to name two). Ceph even has an associated foundation that is a child of the Linux Foundation.
Why is this important?
Aside from the distributed storage capabilities. Perhaps a key factor is that the portability opens up several solution options that can be deployed both to the cloud and on-premises (yes, a POSIX filesystem is portable, but doesn’t offer versioning, and distribution isn’t transparent). One of these is that development around the adoption of HDFS has slowed.
Datalakes and cold-stored data
But S3 isn’t just helping data lakes; it is gaining adoption with OLAP structured storage because it is easy to use with Parquet files and indexes to create partitioning tables. Parquet’s columnar characteristics mean it’s easy to retrieve just the attributes you want (with focus on the rows through the use of indexing and segmentation). This helps power retained data that is cold stored (e.g. historic monthly finance figures kept in a database for analytical and compliance needs).
The open source ColdFront extension for Postgres (led by pgedge) leverages Parquet and S3 to store cold data. Still, the clever thing is that the feature transparently combines the relational data with the cold-stored data in Parquet files stored in S3.
S3 Tables
This confluence of Parquet files to support columnar table data storage has led to the development of a specialization of S3 in the form of S3 Tables. Today, S3 Tables are only an AWS feature, and it will be interesting to see if there is a wider adoption of the specification.
The Oracle database while not explicitly supporting the S3Tables API, can effect the same type of behavior by being able to treat S3 storage as an external table, and handle the metadata that will define a Parquet file within the S3 store as a while or segment of a table (for more read go here).
Ok, as a heading, that is a bit clickbait. But the underlying message is true. Let’s ask ourselves some simple questions. What is a patch? The Cambridge Dictionary says within the context a computers:
The key here is change to make something existing to be correct. That means there is something wrong. this can include a means to circumvent security. But to need a fix, we first need to find the fault. As a result that must mean we there is ALWAYS the chance of a fault. Simply put, there is always a vulnerability before a fix. Therefore regardless of how patched we are there is always a possibility however small of their being a vulnerability and that a malicious actor (or hopefully a bug bonus schema or something like Mythos) finds it first.
Cultural challenge
As an industry, our first question when there is a problem is, have you applied the patches (or bug fixes), or have you upgraded to the latest version yet? This is almost as pervasive as the old joke when there is a computer problem: Have you turned it off and on again?
Don’t get me wrong, if you have patches, I’d err towards applying them. But, I’d also advocate trying make a risk assessment as to what patching could go wrong; if you can safely eliminate a problem, it is better to do so. But, blindly patching can be an issue. Patches can conflict other parts of the system, be applied to systems resulting devices being ‘bricked‘ or hitting the ‘blue screen of death‘.
Let me illustrate, using desktop PC. When I receive a windows update, it takes a couple of minutes to see what is there. If the update looks like a malware signatures update – no hesitation, the risk of not taking the change is bigger than bot. If the update is for a driver, then we’re more cautious, which driver, can I install the older one. Time to make sure there is an OS recovery point. If it’s something like a H1 or H2 cumulative and features update – time to run a backup and a restore point.
In an enterprise software environment, be that monolithic on-premises ERPs, through to vast microservice deployments there are the potential for problems. Particularly when you consider the amount of possible customisation that could be involved (in the Oracle domain you’ll hear about CEMLI). This means the possible permutations makes it impossible for the vendor to provide assurances. Not to mention if you’ve applied Modification level changes a patch may well conflict with your modification.
Security in Depth
But security, and the principle of security in depth, is far more than just patching. But I will always advocate more. The ideal world is that we have layers of security, so if one can’t be patched, then other layers are mitigations, or to use another term often used in enterprise security, a ‘compensating controls’ (the idea that if you can’t address a vulnerability in one place, you have security controls elsewhere that can compensate for a weakness).
But we forget – breaches precede patches
The thing we tend to overlook when we get fixated on the idea that staying patched means we’re secure is that a patch always follows a breach. Remember, somewhere along the line, a ‘hacker’ (preferably a security researcher, penetration tester, or a white-hat hacker looking to earn money through bug bounty schemes) will find a vulnerability and breach a system as a result. It is only now that the exploit (vulnerability) is known that the work on creating a patch to address it.
But sometimes a patch isn’t practical (it has massive performance implications and demands a rewrite of a key element of the solution). Therefore, we can address the issue with mitigating controls – this could be at a code level, e.g. adding upfront checks for a specific scenario, through to preventing a piece of software from being used in a particular way.
A vulnerability doesn’t mean you’re vulnerable
This is the difficult perspective, and makes interpreting vulnerability scanner data hard to understand. Often, vulnerability scanners will look across a server and determine what products and components are deployed, and what versions of that functionality are deployed. Then look up those components for the list of attributed vulnerabilities. This is helpful, you know, you need to give this some due consideration. But we should not forget that software (and particularly enterprise software) has many configuration points that can cause it to follow different paths, and the vulnerability may lie only in one path. The latest version of WebLogic server has 650-700 mBeans (one of the techniques to configure the WLS behavior).
Look at it another way, you’re a single person with a car (which usually has 4 seats). You need the car to commute to and from work, and to do common chores like going shopping. But one of the back seats ended up getting ripped. The garage wants thousands to replace it. This does not stop you from using the car to commute to work and to do other necessary tasks. Yo have probably unconsciously applied a ‘compensating control’ – never take more than 2 other people in the car with you, and prevent a back seat passenger from sitting on the ripped seat. Frustrating – maybe, but not inoperable. Of course, applying that compensating control to use the car as a taxi would not be commercially viable. You, as the car owner, need to do some due diligence, such as ensuring the rip isn’t a symptom of a more fundamental risk? Yes. Do you need to consider the likelihood of needing more than 2 passengers? Yes.
Patch velocity
One of the challenges that are developing is the velocity of patches coming as AI is able to be used to locate potential errors that then get addressed. This creates an issue of the organization’s being able to rollout patches. This may not sound too challenging, until consider
not all organizations have a pure Kubernetes ecosystem with the means to orchestrate node replacement quickly and easily.
Capacity demands (people, and or compute resources) needed for performing verification steps for any change can out strip resources, or impact other change effort – such as legal compliance, or changes to keep you in business (although a breach could be just as catastrophic).
Demand impacts, when organizations are also trying to create bandwidth and/or budget to transition off vulnerable systems can end up being blocked.
Patch volumes for an environment that is already considered ‘fragile’ are going to generate a lot of work, as not only does the patch need to be applied, but a lot of effort will be needed to demonstrate the the wider business the possible impact has been validated.
To combat velocity, is to work on ccompensating controls in other areas of your system. If the application has a networking weakness, consider not patching the software, and ensuring networking configurations block the network port(s) that are associated with the vulnerability. Add monitoring, and checks to ensure the change to the network aren’t accidentally undone. If you do later apply the patch to address the weakness, don’t unpick the network mitigations, as this gives you ‘security in depth’.
Scanning tools
We can also use detection or scanning tooling to help. For solutions that you have development control over, and preproduction environments should be using these tools. But we have to be careful how we interpret the outcomes of these tools. Blindly take their reports can send us chasing issues, that aren’t really issues. While measuring detecting and measuring risk management is a positive thing, it needs to be done in an informed way.
For example, you run the scanning tool with admin level privileges so it can inspect everything. But, then it detects the finger print of an old version of Java with vulnerabilities in the Swing UI library. But the only reason Java is present is to run some CLI admin utilities. It is never used by the core application, there are no deployed Swing solutions in the environment. If we don’t give consideration to the context of the problem, we end up patching an issue that is in all probability going to be a problem.
At the same time, we can use this understanding to prioritise patching. Yes, applying a patch if available is good, but if applying a patch that breaks your tool is not. Spending time patching a vulnerability, when you don’t use the vulnerable code is eating into capacity to actually implement business change which may well eliminate the problem anyway (for example standardising scripting tools on Python).
What does this all mean?
Bottom line is, if you can patch, you have the ability and capacity to do so then it’s better to do so, as it keeps another layer of defence in place. But you don’t have access to patches (out of support etc), capacity to solve everything immediately then intelligent assessment of the vulnerability, risk driven prioritisation, look to mitigation strategies.
Defence in depth is not a security seller’s motto, but a genuine way to ensure that if one point fails, the next should protect you. But if you can’t patch, understand what the vulnerability is, and ensure you have mitigations in place. Ensure the issue and its implications in your context are documented along with the mitigations – so you have auditable content for any audit. You have been told where the minefield is, so you’ve fenced the area off, so people don’t wander into the issue unwittingly.
So we’ve been busy working on our OpAMP solution. We’ve made a number of enhancements since we labelled the code v0.4 back in April. For this post, we’ll look at the features added and where we’re looking next.
There has been a lot of feature development, particularly in support of working with Fluent Bit (and to a degree Fluentd).
Standalone or OpAMP server plugin for editor
We started building the Configuration Editor as a standalone capability. The thinking was that we would refactor it into the OpAMP server once we were happy with the functionality and had progressed far enough. But, as we saw this come together, it occurred to me that both deployments are good, as part of the OpAMP server, seeing the configuration being used is handy, but having a freestanding editor (without worrying about the connectivity to communicate with agents) is also a real use case.
So we have created the setup, where if the OpAMP is told to look for the Editor via the definition of Python endpoints in the configuration and it is deployed, it will be incorporated into the server. If the editor isn’t provided (deployed, or not identified in the configuration it isn’t offered in the navigation.
The editor is completely configuration-driven through JSON, so it only requires extending the JSON to support custom plugins or to enhance the validation rules (e.g., adding a REGEX to how a particular parameter is set). This far outweighs the current Dry run checks.
The configuration is considerable, so we refactored the structure to make it far easier to work with, and created some code that can mine Fluent Bit’s GitHub to generate an initial clean set of JSON docs. Of course these need an eyeballing to ensure they’re correct.
Catalog viewer
The catalogue viewer is a natural extension of the editor and leverages the way we tag configuration files with version details in the editor. The catalogue viewer has a configuration which tells it where to look for candidate files. These are then listed with the metadata.
The catalogue viewer presents all the identified files, which, when selected, if it knows about the config editor, will open the editor with the file. Otherwise, it opens a simple view of the file.
The catalogue viewer works in the same way as the editor in terms of authentication.
CLI
A Command Line tool maybe and odd choice of feature, but we found ourselves creating more and more scripts to support Windows and Bash shells with commonality. So we elected to leverage some frameworks to help reduce the ongoing effort required to maintain them. The utility can be used to generate the command, if you wanted embed the process of starting or stopping a process into the host OS.
Agent as Supervisor or observer
The OpAMP documentation suggests that the agent logic is embedded, or wrapped with a supervisor, with the inference that the agent forks the application process. This means that introducing the OpAMP would be invasive. There is a noninvasive variant of the Supervisor that we’ve called Observer. Here, the agent know how to identify the key process by examining the host’s processes. Then tasks such as restarting require an understanding of how to get the OS to trigger, for example, if the service is known to init.d in Linux, we can use the service command.
Architecture view
MCP and Slack for ChatOps
We’ve got the Slack foundations progressed, so we can use natural language VIA langgraph to then work with MCP, which exposes a subset of API capabilities. We’ve developed both the main server and the Broker to support the MCP endpoints, with the Broker acting as a proxy to the regular Server APIs. This means if we want the MCP to be usable from outside our network, then we can separate the Broker and Agent into separate networks so that the wider set of endpoints the server provides aren’t exposed.
But if you don’t want that the MCP endpoint can be switched on.
What next …
Config deployment
We have started to address the foundations of managing configuration deployment (for example, if our Catalogue Service is used with the Server, it can be used to select the configuration files to deploy.
Part of config deployment is understanding what is deployed, so we have developed some strategies for versioning Fluent Bit and Fluentd, and I think it will translate to other configuration files (at least into the Observability space).
What we haven’t done is fully implement the process. So we’ll focus on that
E2E Testing
There is a lot of functionality here now, including tests with Playwright, but we really need to extend it to provide end-to-end tests in a clean environment that is set up from the various pip and wheel files.
Deployment Artefact access
We also want to start making these artefacts easier to retrieve, such as pulling the Wheel or PIP files from GitHub or PyPi.
I’ve written a bit about AI in the development process; this has been driven largely by my own experiences, colleagues’ experiences, and blog content from people I trust. So I thought it would be worthwhile to validate my perspectives against those who are more in the know on the subject. So here is my review of the book Vibe Engineering by Tomasz Lelek and Artur Skowroński.
The book opens with a very clear differentiation between vibe coding and vibe engineering (which approximates to what I’ve previously called AI-assisted development). Not only are the key conceptual differences outlined, but the consequences of vibe coding into production are also really driven home …
teams that skip the transition from prototype to engineered artifact consistently report higher defect density, longer incident resolution, and faster architectural decay
The book also shares some real horror stories of blindly trusting LLMs, particularly in operational contexts.
The crucial challenges of vibe coding beyond ideation, PoC, and possibly MVP are brilliantly distilled. Code will do something, but is it right? Is it safe? Will it scale? Can we maintain it?
Tomasz and Artur outline a form of debt called trust debt. Where we have trusted the LLM, and it accumulates issues, particularly with NFRs that are not managed and paid down, it will seriously bite, just as tech debt does. The difference is that tech debt is more readily appreciated and generally easier to understand.
debt is a direct byproduct of the dump-and-review culture. This approach uses AI to generate a large slab of code, opens a pull request, and implicitly offloads responsibility for verification to the reviewer. It’s classic diffusion of responsibility: the presence of the AI (“the model wrote it”) and a reviewer (“someone will check it”) dilutes the author’s ownership of quality
Current approaches to this kind of development can very easily lead to the issues that Human-Machine Interface researchers talk about as automation complacency and the out-of-the-loop problem
The book also highlights interesting parallels, such as those in autonomous vehicle accidents. The consequences may not be as spectacular or as tragic (today), but they can be just as harmful, given that code affects every little aspect of our lives and the decisions we make. It is only a matter of time before it is influenced by vibed code. How long before pressure and a failure to comprehend vibe coding vs vibe engineering creeps into mission-critical development?
Once the consequences and challenges are called out, the book takes us on a journey to illustrate how to better approach vibe development, specifically through defining what a successful outcome should be. The brilliantly simple thing here is that the two approaches are demonstrated with multiple different LLMs using the same prompt.
While the book provides brilliantly illustrated proofs for how to better approach vibing (moving from coding to engineering), Tomasz and Artur point out that this alone is not enough; we need to lean into broader process improvements and leverage good engineering practices.
This first chapter then sets everything up that follows, taking on a journey of re-engineering a solution. Illustrating how to prompt to extract from an existing solution the details that can then be fed as prompts to generate a new solution.
The narrative progresses through considerations such as context compression, then leverages tools to enable the LLM to take on significant tasks, such as UI design, by giving it the information it needs to work out and create React components with a consistent look and feel.
The books reveal some really good ideas that allow things to be developed far more efficiently, for example, rather than expecting the LLM to scan through code, exposing the Language Server, which provides a lot of today’s IDE smarts, such as navigation through the call chain in an application. Exposing the LSP as an MCP tool offers the LLM an efficient and more reliable approach to analysing code.
If you want to follow along and test the points that the book makes, you’re not going to need to fork out masses on LLM tokens; the authors are very clear that the cost to repeat the exercises can be done within free/trial service tiers.
Conclusion
I don’t want to spoil your enjoyment of reading the book by revealing its secrets here. But there is a lot of great content, which means that, with some adjustments to how the LLM is prompted and some setup, it becomes possible to significantly reduce that trust debt.
If you’re heading down the road of vibe-based development, I would highly recommend digging into this book. We’re already making some further refinements to our processes. The changes needed to transition from vibe coding to vibe engineering won’t be shocking to those with a software engineering background. But their adoption is likely to pay back significantly.
There are a lot of posts on various platforms about how AI is generating ‘slop’, and it is actually costing a lot of time to take what was generated and put it right. Cleaning up after AI is something people are even incorporating into their online biographies. On the other hand, others indicate they’re getting good results. So what is the reality? I think what we’re seeing is a combination of factors, but there is a healthy dose of human behaviour amplifying the issue.
There is no getting around how well a publicly available foundation AI can perform, depending on the availability of content for training. Providing simple Python logic, which has been asked for with clear precision and expressed clearly, is likely to yield positive results – that’s simply a function of the amount of accessible content on the web. If you asked an AI to generate formal method notations like Z or VDM – good luck.
But when we see what, on the surface, looks good, it is easy to be taken in and start to trust the LLM. Combine that with several other factors:
Humans, by our nature, will tend to minimise effort (or, if you want to be crass with language, lazy). You can see this through things like UX design principles that advocate avoiding ‘cognitive load‘ and ‘choice overload’ (the idea that we can only cope with so much information in working memory, going beyond that, we are more likely to make mistakes or need to apply more cognitive effort), through to George Kingsley Zipf’s Human Behavior and the Principle of Least Effort.
When we hand a task to an AI, the way LLMs work means that it will seek to provide an answer, rather than say sorry, can’t get an answer (if we did that, then the principle of least effort would explain why we’d give up with it). So we’re going to get a result, whereas the non-LLM route means we won’t see its incorrect until we’ve finished. When it comes to coding, the LLM is unlikely to make the coding errors we can make as humans, but the code it produces may not be as elegant or efficient, may not address all the edge cases, and may even miss the problem we want to solve. But the outcome will be executable.
Next issue is the quality of prompting: as humans, we have (generally) good long-term memory, and even if we forget specifics, we build a strong contextual understanding that we use. But the LLM doesn’t have this; it doesn’t know whether we’re trying an idea out or writing code that needs to be bombproof and extremely scalable. We have to define that very explicitly. If we’re working with a new or junior developer, we understand that we shouldn’t make these assumptions and will seek verbal and nonverbal feedback if there are issues with clarity or expectations.
When building things manually, each step we take to create the solution (whether that’s code or a PowerPoint) is slower, and we have more time to evaluate what we want to do, how we want to do it, and why we want to do things a particular way. In some respects, the LLM approach is like code reviewing. For a proper review, the reviewer is going to walk through each line of code and evaluate it – a process that can be lengthy. But under time pressure, what often happens is we’ll take a 1st fast pass to look for the ‘bad smells’ and pick up on any obvious issues. Then zero in on the bad smells, or at least the worst ones, and look more closely. But the time pressure, knowing the tooling we have to help eliminate issues, and the mental effort to quickly understand a lot of code take their toll. This, to varying degrees, is exactly what happens with LLM artefacts, except that the rate at which code can be properly reviewed is now a real issue relative to the rate of generation.
Commerce has always been about either innovating to compete or doing things more quickly and cheaply. That pressure has grown as technology has advanced, creating the potential to do more. As a result, it is not surprising to see that pressure results in code being generated, even if that is likely to drive an unwitting accumulation of technical debt that will bite in the years to come. Furthermore, recognising poor-quality code takes experience. Which means expensive engineers. It is possible to appreciate that a non-technical person can generate a basic desktop utility using ChatGPT.
As you can see, there are plenty of things we, as humans, can do to mitigate ‘AI slop’ and get AI delivering value, quality, and velocity.
Conclusion
The bottom line here is an issue of expectation. We wouldn’t be so harsh as to ask someone with dyscalculia to produce a company’s accounts using only a pencil and paper. Another way to look at it, you’d not ask an unqualified accountant to do your tax return. But that is often what is happening, Someone with dyscalculia could easily ‘hallucinate’ the numbers. An unqualified person is not going to know all the rules needed to complete a tax return well enough to minimise tax exposure.
to err is human; to persist in error is diabolical
Saint Augustine
It is human to make mistakes, and poorly directing (or training) an LLM is certainly an error. But we know this is possible, so we should consider what the code is for and take appropriate steps to mitigate it, working to improve how prompts are given and how context is provided to enable better outcomes.
What is clear to me is that ‘AI slop’ isn’t going away soon, and that as engineers, we have to get better at prompting to get the best code we can from an LLM. While it would be nice to think that the industry will realise that LLMs are not a panacea, and you still need those expensive engineers to prompt an LLM so that they don’t generate unnecessary reams of low-grade, brittle code.
The question really has to be, who is going to build an LLM model and agents that can pre-screen code and call out AI ‘slop’, saving code reviewers (particularly those who are looking after the open-source solutions on which so many of us depend). If Anthropic’s Mythos finds 20-year-old bugs, we should be able to help protect open-source projects from low-quality, poorly prompted AI-generated code.
Something that vendors like Microsoft have been really good at is reducing the friction on getting started – from simplifying installations with MSI files and defaulted options through to very informative error messages in Excel when you’ve got a function slightly wrong. Apple is another good example of this; while no two Android phones are the same, my experience is that setting up an iPhone is just so much easier than setting up an Android phone. It is also the setup/configuration where most friction comes from.
Open-Source Software (OSS), as a generalisation, tend to be a bit weaker at minimising friction – this comes from several factors:
When OSS is part of a business model, vendors can reduce that friction, making their enhanced version more attractive.
OSS contributors are typically focused on the core problem space and are usually close enough to the fine details to not need those fancy features to keep the rest of us out of trouble.
The expectation is that tools to make configuration easy are embedded in the application, making it heavier, when the aim is to keep things as light as possible.
Occasionally, a little bit of intellectual snobbery can creep in
The common challenge
The issue that I have observed is that we often go through cycles of working with a technology. For example, you’re building a microservice. Chances are, you’ll start writing and running it locally, without worrying about containerization. Once you’re pretty happy with things, you’ll Dockerize the service, start testing it locally, and then you’ll be ready to deploy it to a cluster. Now you’ll need your YAML. It may well be weeks since you last looked at Helm charts. You end up cutting and pasting your last configuration. But now you need to use another feature of Helm, can you remember the exact settings for the feature. So now you’re trawling the net for documentation, and then it takes several tries to get it right.
AI may well step in to help developers in this area, where solutions and products are well-documented. But with the wrong model or insufficient detail in the prompt, it’s easy to make a mistake. Personally, I’d turn to AI when it becomes necessary to trawl code to better understand the configuration and its behaviour, and to set options.
Experimental Solution
Solution – well, that depends upon the configuration syntax. We have been experimenting with RJSF (React JSON Schema Form), which provides a React-based UI that can be dynamically driven by a JSON schema and validate data with AJV (an alternative stack considered would have been around JSON Forms).
{
"type":"object",
"title":"Dummy",
"properties":{
"name":{
"type":"string",
"const":"dummy",
"title":"Plugin"
},
"copies":{
"type":"integer",
"description":"Number of messages to generate each time messages are generated.",
The above fragment shows part of the Schema definition for the Dummy plugin for Fluent Bit.
By then creating a schema that defines the different plugins, attributes, etc., we can drive validation and menu items easily in the UI. Admittedly, the config file is significant given all the plugins and configuration options, but it is a fair price to pay for a UI that validates the data. Establishing the schema to start with, we’ve covered it through scripting the retrieval and scraping of the Fluent Bit pages, which are pretty consistent in structure.
We have added some custom elements into the definition, for example, x-doc-reference, which allows us to extend the React components to provide features such as a link back to the original documentation as you select attributes or plugins.
As a result, we very quickly have a UI that can look like this:
A lot easier to view and tweak, with no need to hunt for valid options. Even if we want more information, we’re just a button click away from the open-source data. Perhaps we should provide a version that hyperlinks to the Manning Live Books on Fluent Bit, etc.
There are a few other factors to consider; for example, Fluent Bit configuration is YAML, not JSON, which can be easily resolved given the relationship between the two standards. Then there are processors that can embed Lua code or a SQL-like syntax. As we’ve chosen to provide a Python backend, we’ve addressed this by providing REST endpoints which can query out of the JSON the code or SQL and perform validation using the Python Lua Parser, and the SQL syntax can be addressed using the Lark library for processing the SQL, as the syntax is simple enough to define and maintain the syntax.
Outstanding Gaps for Fluent Bit
We still need to address several features that Fluent Bit has, specifically:
Environment variables
Includes
These issues should be straightforward to overcome, although dynamically including the included elements into the UI view elements can be done. The challenge is: if any changes need to go into something that has been included, how do we push them back to the included file? Particularly if there are multiple layers of inclusion.
What about Fluentd?
Fluentd configuration isn’t JSON-based notation, but it is structured. So, to apply the same mechanism, we’ll need to define a schema and a mapping mechanism. The tricky part of the schema is that Fluentd supports nesting plugins, since the way pipelines are defined for routing differs. While JSON schema will enable this with constructs such as anyOf, oneOf, object nesting, and bounded object arrays, the structure will be more complex.
The second challenge will be the transformer/renderer, so we don’t introduce issues from having to escape and unescape characters, since JSON Schema is stricter about character use.
Then What?
Well, if we get this going, we’ll probably incorporate the capability into our OpAMP project and maybe create a build that lets the configuration tool run independently. Lastly, perhaps we should look to see if we can make the different layers a little more abstract, so we can plug in editors for other configurations, such as OTel Collectors or the ELK Stack.
As a bonus, perhaps transform the Schema into a quick reference web document?
You must be logged in to post a comment.