Gen AI and its Impact for CISOs and Their Teams in 2024 with Liqian Lim at AIE 2024
In 2023, we saw the explosion of AI, with modern coding tools such as GitHub Copilot, AWS CodeWhisperer, and ChatGPT coming on to the market. All of these generative AI coding assistants brought AI to the forefront of the news, and propelled AI into a top security priority. In late October 2023, the White House’s new executive order around AI was a timely reminder to organizations that security should be a key concern for any organization leveraging AI.
Transcript
Hello and welcome to Jen AI and its impact for CISOs and their teams in 2024. I'm Lim Lichen or rather Lichen li, senior product marketing manager for AI at Snyk. Briefly, I spent 10 years in law before doing my MBA and transitioning out into tech.
I spent the next six years in various functions across AI and blockchain, so I understand risk management, AI for enterprise and strategy. Now this is what we'll be covering today. First we'll talk about the landscape.
Then we'll cover three areas for building trust, firstly through AI safeguards. Next through fine tune hybrid models, and then thirdly through security expertise. And then we'll wrap up with takeaways.
Now, I'm sure you've seen the kinds of reactions I have when the conversation turns to ai. You get a camp that screen Skype and runs for the HA and you get a camp that adopts AI with gusto just like my parents with WhatsApp spam messages. And then you get the in-betweeners like sny who understand what that AI is just the tool.
And like any other tool, it isn't inherently good or bad. There's a balancing act involved in getting the best out of it. I want to stress that the AI landscape is so new and moving so quickly and no one has all the answers.
We could benefit a lot from pooling our experiences and learning together, which is why today is about sharing sneak discoveries and experience as pioneers and Gartner and forester leaders in the A area of AI powered application security so that you can see principles put into real action. Trailblazing can be lonely, but we think that a market has caught up and we're proud of the fact that we've been doing some of the key things that Gartner's now recommending well before they release that report. Title Four Ways generative AI will impact CSOs and their teams.
I'll be referencing this report from time to time and I'll just call it the Gartner report. When I do, we will make copies of this report available to you. There are many categories of ai.
The main ones I'll talk about today are the AI big umbrella, two main categories, the symbolic AI category and machine learning category. And under machine learning it under machine learning is a subset of generative AI or gen ai and branching out from that large language model or LLM. There are many other categories, but I'm only going to focus on these four today.
I remember asking my expert colleagues to speak English in my early days in AI when they threw a bunch of terms at me. So before we dive in, I'm going to go over some terminology quickly and I really mean quickly. So symbolic AI is a type of AI that processes knowledge represented in the form of symbols with rules and logic.
How this works is that human experts manually encode knowledge and rules into the system defining the symbolic representations and relationships. Importantly, in unlike machine learning, symbolic AI uses formal logic and inference rules where such rules dictate how the system should process and manipulate symbols to reach conclusions or make decisions to derive new knowledge from existing symbols. And this strict adherence to rules and knowledge base is why symbolic AI is used to develop expert systems.
Now, machine learning is a subset of AI that enables a system to automatically learn patterns from data sets and then use this data to refinance performance like making predictions or performing specific tasks, importantly without explicit programming. Then we have generative AI or gene ai, which is a subset of machine learning ai. It creates new content like images, audio and text that resembles or fits within a given dataset rather than simply analyzing existing data.
Generative models learn patterns from training data and then used as knowledge to generate novel output, which is why quality and type of training data, for example, using permissively licensed training data to avoid copyright issues is important in maintaining quality and usability of outputs. And because of the way that this technology works, hallucinations in errors are inherent in its outputs, but this can be managed. We'll talk more about this later.
A large language MO language model or LLM is a subset of generative AI and therefore a subset of machine learning AI that is a type of artificial neural network pre-trained on massive text-based data sets, allowing it to perform different natural language processing or NLP tasks such as recognizing, translating, predicting or generating human-like text or other content and to understand language to some extent, you already know by now that due to the nature of gen ai, there are risks inherent within it because of one, the very nature of how this technology works, learning patterns, creating its own logic from those patterns and generating outputs based on that logic. And number two, the fact that gen AI output is only ever as good as the data that goes into training it. So there are issues with biases, errors and hallucinations when you work with gen AI outputs.
Now we talk, we are moving into talking about the uh, landscape adoption and concerns. So we have lived through general generational waves of technological evolution, but we're, what we're facing with AI is unprecedented. We are already clocking in at two times the speed where internet adoption was in its early phases.
Not only is there unprecedented speed of adoption, but Gartner also predicts massive scale of adoption forecasting that over 80% of enterprises will have used some form of AI or deploy their own AI model by 2026. So with this seismic shift comes urgency of action. CSOs are aware that this amazing technological revolution has already started happening and although they want to get on board, they have some pretty big concerns around generative ai.
According to Gartner, some of the most common areas of concern for CSOs around gen AI include copyright violations, biased, incomplete or erroneous responses, policy violations and lack of transparency. At snyk, we've also heard this a lot and although CISOs are aware of the urgency of these issues, there are a multitude of security solutions out there all claiming to do the same thing. But how do you know which one's right for your organization and where to start?
All these lead back to trust. You need to be able to trust whatever you consume from your AI to innovate freely. In order to trust your AI outputs, you need to be confident that it's secure.
The area of AI is fast, but today we'll be focusing on gen AI and software development. And I'm going to try and simplify things by zooming into the actions you can take to increase trust in your AI generator code. Reduce the mental load of adopting a gen AI tools and take back control whilst enjoying the benefits of ai.
This will be a three-pronged approach, safeguards hybrid, fine tuned models and security expertise. Trust through AI safeguards. AI is an evolving area, so everything is new and everyone's still figuring out what safeguards to put in place.
However, the approach I'm going to talk about isn't new. It's still based on foundational principles that we already have in securing software development. First people, humans are complex creatures.
Does anyone ever take the medicine when a doctor doesn't explain why you've got to take it? The answer is no. So there's limited impact in having tools and processes without empowering your people to understand why they need these and how to use these tools correctly.
Gardner highlights training as a key component to mitigating gen AI risks and we agree to manage gen AI risks. You'll need to provide training and continuous education. For example, interactive bite-sized sessions, sessions and lessons like those found and sneak learn or embedding.
Learn to go training in your tools like having inline advice on the vulnerabilities found by a security tool. Two processes. These should not be occasional checkbox exercises because they don't stick and there's a chance of slipping up one day.
You want long-term habits as part of your AI governance. Use simple pithy policies and procedures. For example, appointing AI champions, setting up permitted vendors usage guidelines, et cetera to help guide your teams in the short term.
But for long-term genuine adoption, you need to align your policies and procedures with existing workflows to get buy-in and collaborate with the security and developer teams to, to create and improve these policies. It iteratively, some of the mistakes we've seen are lengthy long-winded, uh, policies and procedures that gather dust and that no one reads or follows. And another big mistake is creating, uh, policies and procedures that don't make sense for certain functions they impact on because not all stakeholders were consulted in a process and therefore the policies were not formed collaboratively.
And lastly, tools enable your team with the right tools for the job. There's no point in having AI fast code but security that lags behind or doesn't get adopted because it disrupts user workflows. You should never adopt an AI coding assistant without a suitable security companion in place.
I'll go into more detail later. For the safe use of ai, you need multiple layers of safeguards in order to avoid any single point of failure. Like Gartner, we've found that there's a strong push for progress whether or not security is ready.
Sneaks 2023, AI generator code security report shows that 80% of developers are bypassing security policies to leverage the power of AI coding assistance. And the Gartner report has simil similarly found that 89% of business technologists would bypass cybersecurity guidance in order to meet a business objective. And they, they go on to say that organizations should expect gen shadow generative ai, especially from business teams whose main task is to generate content or code.
The long and short of it is that people, processes and tools all need to be in place to work together to effectively set the stage for trusting your AI generator code. This takes, this takes us around to fine tune hybrid models and Garner says that businesses planning to use AI should adapt to hybrid development models such as custom gen AI applications with in-house model design and fine tuning. And this is something that SNY has been doing for a few years now, but why AI at all and why go to the trouble of having a proprietary hybrid AI model?
The answer to the first part of the question is simple speed and scale faster than developers, AI power coding assistance at a new left and you need a security solution that's at least as fast as these assistance. Also given volume of code created by AI assisted developers and coming proliferation of new vulnerabilities and new methods of attack, you'll need a security solution that can learn and scale rapidly. For the second part of the question, why a hybrid AI model?
These are the same reasons why businesses have reservations around trusting AI generated outputs and automating actions with these outputs because you don't know what's going to come out of your generative AI and you don't know how to trace its processes due to varying levels of capacity in the workings of gen AI models. Let me give you a couple of easy metaphors for this. You've already learned at the start of this webinar that symbolic AI is strictly based.
It's like the school head mistress who adhere strictly to rules and a knowledge base. Generative ai, however, is the dreamer. It comes up with a BA bunch of different things, but you're not sure which ones are product of its imagination.
You can think of this as coherent nonsense and which manufactured things are factual, strict rules and a curated knowledge base can help to reduce uncertainty around provenance and reliability of output. This is why we've tethered the machine learning dreamer elements of our AI machine to a symbolic AI school, school head MRTS element and why we advocate for the same. Let's talk a little bit about our in-house developed hybrid, uh, about how our in-house developed hybrid model works.
So you understand why we and Gartner believe this approach to be best practice, Snyk has a purpose-built solution dedicated to only one thing code security and the integrated approach of combining symbolic AI with machine learning gives the best results through engaging the right methodology at the right time. Combining symbolic AI machine learning and LLM models all focused, all focused on one single developer security workflow enables results that are unattainable with a single model. In this example, symbolic AI is used to enable us to test for issues based on an abstracted understanding of the data flow.
This makes for more accurate results because you are testing for flows, sources, sanitizer, sinks rather than string matches. To illustrate why falling data flows is so important, we can use a pinball table analogy. Every line snippet block of code changes the way that your data flows.
So imagine the data as a pinball on the pinball table and every new bit of code as a bumper on your pinball table. Every time a new bumper is added, it changes the direction that the pinball takes and changes the pinballs interaction with the rest of the table. So you don't know where the new pitfalls are.
This is why analyzing the entire application or table and following the data flows results in fewer false positives and is crucial for the integrity and accuracy of code security. Next, the machine learning element. We use machine learning to generate new codes that we can apply to the searches, increasing accuracy and coverage over time.
When a vulnerability is found, we use an LLM to generate a fix, but then to double check that accuracy of the LM fix and remove the risk of errors or hallucinations, we put a version of the fixed code through a test before even showing the proposed fix to the user to ensure that a fix works and doesn't create new problems in the rest of the code. And throughout, because AI doesn't have cognition, we have human experts curating training data, tweaking inputs and checking and refining outputs. This hybrid model with human fine tuning reduces gen AI errors and hallucinations and significantly increases accuracy, yet reviews applications completely.
As proof of this method's effectiveness, we have an oasp benchmark accuracy of 73% compared to the 54% accuracy of a much beloved developer brand. And this level of accuracy, it's only increasing and growing with time. If you only coupled symbolic AI with the LLM fix alone or just a machine learning element, you could not provide the real time solution of a fast proactive yet accurate and complete scan in the IDE with issues detected and fixes applied all in one flow using only one of the LLM or machine learning would either require the developer to repeat steps killing the productivity gains from using AI coding tools or potentially result in the application of fixes that cause insecurities within your code base.
Finally, and importantly, our multimodal approach ensures that we don't regress to the mean despite crowdsourcing our data sets. When more general purpose LMS learn, they need to draw from a massive and diverse dataset, which includes a large variance of coding standards. This logically reduces the average standard for the output of such rrms.
However, if you have a purpose-built security tool with human experts that constantly in a loop, you can more reliably maintain state-of-the-art standards and boost precision in your ai. The takes us through to security expertise. The Gartner's report reports, uh, the Gartner reports recommendation that businesses should choose fine tuned or specialized models that align with the relevant security use case for more advanced teams is exactly what we've been doing in advocating for security is a highly complex area and code quality is a different beast from code security.
It's not easy to spot security issues yet. Security is important. Not only is our AI model customized for security, it is as mentioned above proprietary and we remain independent of AI coding assistant brands.
What does this mean? Well, not all security tools create fixes, but of those that do, some have been built by the same brands that built AI coding assistance as well. This means that the fixes rely on the same AI models behind these coding assistant.
These AI models or LLMs were trained for code functionality, not security. Given how good AI's outputs are, depends generally well depends heavily on the quality, context and accuracy of its inputs. It's arguable that an LLM creates it and trains specifically to secure code will produce fixes that are more reliable than those coming out of a more general purpose.
LLM having constant and in-depth input from security experts result in an AI powered tool that's laser focused on strategic security. And this means making streamlined findings and impactful fixes as opposed to having exponentially growing vulnerabilities of every level of priority thrown at your teams. Offering custom rules to further tailor outcomes to your own policies and presenting centralized cross team security reports so people can focus on fixing the issues most impactful for and unique to the business.
Garner also goes on to state that generative cybersecurity AI on real-time content may take more time to arrive as it is likely to require specialized and maybe smaller models trained on such data. We couldn't agree more with our tailored AI model trained on specific data. In fact, we've already been providing generative cybersecurity AI on realtime code ahead of the report because of our specialized model.
Whether a security provider creates its own AI model, how it creates its model, when and when and how frequently human security specialists are involved all contribute to how robust a security tool is. Our technology and process came out of our collective knowledge and the letter is ultimately our biggest differentiator in the report. Gartner also anticipates that if generative AI becomes increasingly more proficient at uncovering new vulnerabilities in code zero day attacks, which basically mean attacks exploiting undisclosed vulnerabilities will become more common and that this will accelerate the development of more supply chain attacks against widely deployed and privileged applications.
This will likely be an area where rapid react reactive responses necessary. Thus, and user organizations should evaluate tools and services to monitor the software dependencies so it is not weather, but when generative AI will start speeding up zero day attack occurrences, therefore a security tool that not only proactively neutralizes risks in the IDE before they proliferate in your pipeline, but also continues to protect you across the SDLC with native integrations and a consistent comprehensive overview of your security program will give your organization a competitive advantage. For example, sneak open source allows for easy generation of SBOs and usage of the data within for insights to defend against unknown risks.
Like zero day vulnerabilities is tricky, but you can manage the degree of chaos you encounter when the unexpected happens. By planning and with structure sneak app risk helps security leaders with this, but generally look for a solution that gives an overview of your security posture and program shows you where the gaps are issue of fixed trends and helps leaders with strategy creating and with planning for an uncertain future. Your chosen solu solution should have deep senior level security expertise behind it as they will understand what to prioritize in a consolidated consistent view.
No solution can help you to cover every eventuality, but you can be better prepared for the unexpected. And finally, takeaways. The speed and scale of AI adoption is unmatched, so you need to act fast to innovate safely.
And we have three recommendations for how to um, take you closer to innovating with trust and safely. Firstly, layers of safeguards to hedge against risk, you need people, processes, and tools. Secondly, you need to look out for proprietary fine tuned hybrid models that are specific to your needs.
And thirdly, look out for security expertise because this will make a difference to the quality of your output. It will bring you impacted targeted results and also give you a better user experience overall. And finally, trust your safeguards with this multi fronted approach and you can begin to trust your AI generated code.
This was brought to you by snyk, trusted security at the speed of ai. If you have any questions, you can forward these to sync and include your email address. We'll make sure that we get back to you.
Thank you very much for your attention and time.