It is funny how when websites reference S3 (Simple Storage Service) storage, they typically reference a couple of the Hyperscalers (usually from Azure, GCP and AWS). While I get the reference to AWS, as they developed the service that has become the foundation for the API specification. But the association with just those hyperscalers is deeply misleading (and, to be honest, it bugs me). Any infrastructure as a service (IaaS) provider provides an S3 solution, and would probably not be taken seriously if they didn’t.
But S3 is not just a cloud-only capability today; enterprise-class SANs have S3 as part of their offering. There are open-source offerings that will allow you to overlay S3 onto commodity hardware Ceph, MinIO, to name two). Ceph even has an associated foundation that is a child of the Linux Foundation.
Why is this important?
Aside from the distributed storage capabilities. Perhaps a key factor is that the portability opens up several solution options that can be deployed both to the cloud and on-premises (yes, a POSIX filesystem is portable, but doesn’t offer versioning, and distribution isn’t transparent). One of these is that development around the adoption of HDFS has slowed.
Datalakes and cold-stored data
But S3 isn’t just helping data lakes; it is gaining adoption with OLAP structured storage because it is easy to use with Parquet files and indexes to create partitioning tables. Parquet’s columnar characteristics mean it’s easy to retrieve just the attributes you want (with focus on the rows through the use of indexing and segmentation). This helps power retained data that is cold stored (e.g. historic monthly finance figures kept in a database for analytical and compliance needs).
The open source ColdFront extension for Postgres (led by pgedge) leverages Parquet and S3 to store cold data. Still, the clever thing is that the feature transparently combines the relational data with the cold-stored data in Parquet files stored in S3.
S3 Tables
This confluence of Parquet files to support columnar table data storage has led to the development of a specialization of S3 in the form of S3 Tables. Today, S3 Tables are only an AWS feature, and it will be interesting to see if there is a wider adoption of the specification.
The Oracle database while not explicitly supporting the S3Tables API, can effect the same type of behavior by being able to treat S3 storage as an external table, and handle the metadata that will define a Parquet file within the S3 store as a while or segment of a table (for more read go here).
A lot of developers are focused on CVEs, as we need to know does the library, language or other building block I use to create a solution have an identified vulnerability that needs mitigating? We probably also pay attention to OWASP (Open Worldwide Application Security Project). But CVEs and OWASP aren’t the beginning and end of the story. This a larger ecosystem of standards involved.
CWE (Common Weakness Enumeration) is of particular interest if you get a chance to step back from fire fighting CVEs. This is interesting because, rather than addressing individually identified vulnerabilities, like the OWASP Top 10s, it looks at a classifies the vulnerabilities. When you get down the Base and Variant categories we’re into specific details. This is helpful, as classifying our CVEs against CWE can tell us the technical domains where the maximum effort can return the greatest level of mitigation/remediation of CVEs (be that creating a patch or defining a mitigation strategy).
I was going to sketch out the relationships between the key standards, but realized an LLM can do a better job, so here is a visual summary:
With so many TLAs in the diagram, here is a quick reference pulled together (including from my own library of handy links).
CWSS is a CWE-related scoring method for prioritising software weaknesses, but the older version referenced by MITRE is marked obsolete. (Common Weakness Enumeration)
Scores weaknesses, not CVE vulnerabilities; largely secondary compared with CVSS in current vulnerability workflows.
CVSS v4.0 is the current major CVSS standard family used to express vulnerability severity with base, threat, environmental and supplemental metrics. (FIRST Forum) (Forum of Incident Response and Security Teams)
Maintained by FIRST; it scores vulnerabilities but does not allocate vulnerability IDs. Defines the scoring structure/ranking scale, not ID allocation.
CVD is the process of coordinating information among finders, vendors and other stakeholders before public disclosure and mitigation communication. (CERT Coordination Center)
Process framework around disclosure; not itself an ID system.
SWID tags are ISO/IEC 19770-2 software identification files that can support software asset, patch and vulnerability management. (NIST Computer Security Resource Center)
Product identification standard related to CPE-style inventory matching.
OSV provides a machine-readable vulnerability format and database that maps open-source vulnerabilities precisely to package versions or commit hashes. (OSV)
Adjacent vulnerability-data ecosystem, especially useful for open-source dependency scanning.
Over the years we’ve created a number of GitHub repos with tools and associated content that can help us in the Observability space, such as log generation and replay, to help validation Fluentd and Fluent Bit configurations (or an OTel collector of choice). This extends to the most recent efforts with imp,eating the OpAMP capabilities and configuration UI.
Tools are good, but documentation that is easy to work with is a key to helping users be successful. To support that we have established a consolidated set of pages using GitBook which can be found at:
The nice thing about GitBook, is that it brings together the various resources such as markdown , HTML etc and creates a friendlier experience. It also keeps in sync with the GitHub repos.
New logos
As we’ve established the GitBook content, and using AI makes us less artistically able people to create graphics and logos, we’ve created a couple of logos for the older repos as well:
Fluent Bit Classic configuration to YAML conversion tool.
Log Generator can create synthic logs, or replay application logs to simulator a real system logging.
Ok, as a heading, that is a bit clickbait. But the underlying message is true. Let’s ask ourselves some simple questions. What is a patch? The Cambridge Dictionary says within the context a computers:
The key here is change to make something existing to be correct. That means there is something wrong. this can include a means to circumvent security. But to need a fix, we first need to find the fault. As a result that must mean we there is ALWAYS the chance of a fault. Simply put, there is always a vulnerability before a fix. Therefore regardless of how patched we are there is always a possibility however small of their being a vulnerability and that a malicious actor (or hopefully a bug bonus schema or something like Mythos) finds it first.
Cultural challenge
As an industry, our first question when there is a problem is, have you applied the patches (or bug fixes), or have you upgraded to the latest version yet? This is almost as pervasive as the old joke when there is a computer problem: Have you turned it off and on again?
Don’t get me wrong, if you have patches, I’d err towards applying them. But, I’d also advocate trying make a risk assessment as to what patching could go wrong; if you can safely eliminate a problem, it is better to do so. But, blindly patching can be an issue. Patches can conflict other parts of the system, be applied to systems resulting devices being ‘bricked‘ or hitting the ‘blue screen of death‘.
Let me illustrate, using desktop PC. When I receive a windows update, it takes a couple of minutes to see what is there. If the update looks like a malware signatures update – no hesitation, the risk of not taking the change is bigger than bot. If the update is for a driver, then we’re more cautious, which driver, can I install the older one. Time to make sure there is an OS recovery point. If it’s something like a H1 or H2 cumulative and features update – time to run a backup and a restore point.
In an enterprise software environment, be that monolithic on-premises ERPs, through to vast microservice deployments there are the potential for problems. Particularly when you consider the amount of possible customisation that could be involved (in the Oracle domain you’ll hear about CEMLI). This means the possible permutations makes it impossible for the vendor to provide assurances. Not to mention if you’ve applied Modification level changes a patch may well conflict with your modification.
Security in Depth
But security, and the principle of security in depth, is far more than just patching. But I will always advocate more. The ideal world is that we have layers of security, so if one can’t be patched, then other layers are mitigations, or to use another term often used in enterprise security, a ‘compensating controls’ (the idea that if you can’t address a vulnerability in one place, you have security controls elsewhere that can compensate for a weakness).
But we forget – breaches precede patches
The thing we tend to overlook when we get fixated on the idea that staying patched means we’re secure is that a patch always follows a breach. Remember, somewhere along the line, a ‘hacker’ (preferably a security researcher, penetration tester, or a white-hat hacker looking to earn money through bug bounty schemes) will find a vulnerability and breach a system as a result. It is only now that the exploit (vulnerability) is known that the work on creating a patch to address it.
But sometimes a patch isn’t practical (it has massive performance implications and demands a rewrite of a key element of the solution). Therefore, we can address the issue with mitigating controls – this could be at a code level, e.g. adding upfront checks for a specific scenario, through to preventing a piece of software from being used in a particular way.
A vulnerability doesn’t mean you’re vulnerable
This is the difficult perspective, and makes interpreting vulnerability scanner data hard to understand. Often, vulnerability scanners will look across a server and determine what products and components are deployed, and what versions of that functionality are deployed. Then look up those components for the list of attributed vulnerabilities. This is helpful, you know, you need to give this some due consideration. But we should not forget that software (and particularly enterprise software) has many configuration points that can cause it to follow different paths, and the vulnerability may lie only in one path. The latest version of WebLogic server has 650-700 mBeans (one of the techniques to configure the WLS behavior).
Look at it another way, you’re a single person with a car (which usually has 4 seats). You need the car to commute to and from work, and to do common chores like going shopping. But one of the back seats ended up getting ripped. The garage wants thousands to replace it. This does not stop you from using the car to commute to work and to do other necessary tasks. Yo have probably unconsciously applied a ‘compensating control’ – never take more than 2 other people in the car with you, and prevent a back seat passenger from sitting on the ripped seat. Frustrating – maybe, but not inoperable. Of course, applying that compensating control to use the car as a taxi would not be commercially viable. You, as the car owner, need to do some due diligence, such as ensuring the rip isn’t a symptom of a more fundamental risk? Yes. Do you need to consider the likelihood of needing more than 2 passengers? Yes.
Patch velocity
One of the challenges that are developing is the velocity of patches coming as AI is able to be used to locate potential errors that then get addressed. This creates an issue of the organization’s being able to rollout patches. This may not sound too challenging, until consider
not all organizations have a pure Kubernetes ecosystem with the means to orchestrate node replacement quickly and easily.
Capacity demands (people, and or compute resources) needed for performing verification steps for any change can out strip resources, or impact other change effort – such as legal compliance, or changes to keep you in business (although a breach could be just as catastrophic).
Demand impacts, when organizations are also trying to create bandwidth and/or budget to transition off vulnerable systems can end up being blocked.
Patch volumes for an environment that is already considered ‘fragile’ are going to generate a lot of work, as not only does the patch need to be applied, but a lot of effort will be needed to demonstrate the the wider business the possible impact has been validated.
To combat velocity, is to work on ccompensating controls in other areas of your system. If the application has a networking weakness, consider not patching the software, and ensuring networking configurations block the network port(s) that are associated with the vulnerability. Add monitoring, and checks to ensure the change to the network aren’t accidentally undone. If you do later apply the patch to address the weakness, don’t unpick the network mitigations, as this gives you ‘security in depth’.
Scanning tools
We can also use detection or scanning tooling to help. For solutions that you have development control over, and preproduction environments should be using these tools. But we have to be careful how we interpret the outcomes of these tools. Blindly take their reports can send us chasing issues, that aren’t really issues. While measuring detecting and measuring risk management is a positive thing, it needs to be done in an informed way.
For example, you run the scanning tool with admin level privileges so it can inspect everything. But, then it detects the finger print of an old version of Java with vulnerabilities in the Swing UI library. But the only reason Java is present is to run some CLI admin utilities. It is never used by the core application, there are no deployed Swing solutions in the environment. If we don’t give consideration to the context of the problem, we end up patching an issue that is in all probability going to be a problem.
At the same time, we can use this understanding to prioritise patching. Yes, applying a patch if available is good, but if applying a patch that breaks your tool is not. Spending time patching a vulnerability, when you don’t use the vulnerable code is eating into capacity to actually implement business change which may well eliminate the problem anyway (for example standardising scripting tools on Python).
What does this all mean?
Bottom line is, if you can patch, you have the ability and capacity to do so then it’s better to do so, as it keeps another layer of defence in place. But you don’t have access to patches (out of support etc), capacity to solve everything immediately then intelligent assessment of the vulnerability, risk driven prioritisation, look to mitigation strategies.
Defence in depth is not a security seller’s motto, but a genuine way to ensure that if one point fails, the next should protect you. But if you can’t patch, understand what the vulnerability is, and ensure you have mitigations in place. Ensure the issue and its implications in your context are documented along with the mitigations – so you have auditable content for any audit. You have been told where the minefield is, so you’ve fenced the area off, so people don’t wander into the issue unwittingly.
You must be logged in to post a comment.