top of page

OffSec AI Red Teamer (OSAI) Course and Exam Review

Introduction

For a bit of context, I am a Cybersecurity Expert with extensive experience in Penetration Testing and Red Teaming, I have several certifications including the OSCP, OSEP, OSWE, OSWP, CRTO and others. I have also dabbled with bug bounty, I presented my research at several conferences, and I have reported 12 CVEs. The aim of this introduction is not to brag, but rather to clarify my level of experience. I don't want to set wrong expectations by saying that it was hard or easy without providing sufficient context about my experience.

The Course

I enrolled in the course almost as soon as it was published. My initial plan was to do the OSED in order to get the OSCE3, however, considering how widespread AI had become, I thought it would be interesting to learn more about attacking these systems. To be fair, I was skeptical initially for several reasons:

  1. The course content was new: Typically it takes a few iterations before a course (and the labs) reach their best version.

  2. Information was scarce: Since at the time there was no information about the exam, nobody really knew what to expect.

  3. I was biased: I thought that it will all be about jailbreaks - I was ignorant related to most of the types of attacks against AI.

This being said, I started the course and advanced through a few chapters. It was not bad. Actually, it was quite fun. Some challanges where a bit finicky, since there are LLMs behind, reviewing the responses, so due to the nondeterministic nature of LLMs, you could submit an answer one day and it would show up as wrong, but another day, it could show up as correct. For this reason, I stopped actually doing most of the labs and focused on the course materials. This went on for about a week, but then I had a very busy period so I left the OSAI aside for a while. Nevertheless, the course materials seemed really good - easy to read, high quality! I am not going to go into the syllabus, since you can find it here: https://www.offsec.com/courses/ai-300/.

The Exam

I was able to schedule my exam for the 9th of August at 0900. However, at the time all I knew was that it was going to be a 24-hour exam with another 24-hours for reporting with all the standard OffSec procedures (proctoring, document verification, etc.). Since the earliest OSAI exam date was July 15, 2025, there was barely any information online. However, within a few weeks before my exam, I started noticing several people posting their experience with the OSAI on LinkedIn, so I started collecting a bit more information about the exam. The most important (and surprising) aspect for me was that students are allowed (and encouraged) to use AI during the exam. This was fantastic, even more so considering that my application for Anthropic's Cyber Verification Program (CVP) got approved!

So with this in mind, I felt a lot more confident about the exam. At 8am on the day of the exam, I started setting up my workstation. I have been using AI for various pentesting activities and even tried doing a few boxes on HackTheBox with AI, so i was quite familiar with guiding AI through attacks, but I haven't really set up anything specifically for the OSAI.

At 9am I started. The exam was responsive - absolutely no connectivity problems (the connection is done via Tailscale rather than the classic OpenVPN).

All the context I was given at the beginning of the exam was the following:

  • There are two external targets

  • Obtaining access to any of the two external targets gives access to the internal network

  • The internal network consists of several hosts

  • There are flags obtained from AI vectors worth 15 points, flags obtained from traditional vectors worth 10 points, and a flag from the domain controller worth 5 points

  • Not all machines are vulnerable

  • You must score 75 points to pass

  • Also, unlike other OffSec exams, an interactive shell is not required to complete OSAI exam objectives, meaning that in theory you could only leverage a file-read to get the flag and that would be enough

So I ran a quick nmap scan against the two targets and started some manual recon. I passed my recon notes to Claude Code and by 10:55, I already had my first 15 point flag (AI vector). From the internal network, I decided that there is no point in me running commands so I just guided Claude. I tried to switch to Codex at times, but Codex blocked most of my actions. However, at 11:55 I had my second flag, this time a 10-point one.

With deep access into the internal network and some specific Windows attacks, Claude started blocking me despite being approved for the Cyber Verificaion Program. My solution was to clear the context and start fresh. On each fresh start, I would give Claude the context and ask it to make sure that we are in the exam environment. This way, I found that it took a bit longer until I was blocked again.

Around the mid-day point, I realized that instead of having several Claude Code sessions in parallel, it was more efficient to ask Opus to think of test cases and delegate them to several Sonnet sub-agents. This way, the more powerful (and costly) model would do the reasoning before passing execution to cheaper sub-agents. Moreover, this approach would prevent duplication of the work.

At 19:34, roughly 10 hours after starting the exam I reached 70 points. I was super excited, thinking that I will easily get the remaining points. This was not the case. In fact, I was just going around in circles. I felt like I exhausted all possible test cases. Around 23, I decided to go to sleep and set an alarm for 5am. This would give me another 3-4 hours to finish the exam. Of course, before going to sleep, I just told Claude that I am going to be away for a few hours and to dig as deep as possible. While I fell asleep pretty quickly, at 1am I was already up. So I went back to the laptop. While I was asleep, there wasn't any significant progress, so I decided to make a list of everything that I had up to that point and the remaining machines. Here is where my intution came into play and I just felt that one of the machines must be the one. So with a bit of guidance from me, by 2am Claude got me enough points to pass. Since an integral part of the exam is the report, I wanted to make sure I had all necessary evidence. For this, I asked Claude to create scripts for getting every single flag one by one. This way I could run them manually and take screenshots. Once I had the screenshots, scripts and storyline in mind, I asked Claude to create the report. It was good. So before 4am I already submitted my report, closed the exam, and went back to sleep. Two days later, I got an email saying that I passed:

  • August 9 0900 - Started the exam

  • August 10 0400 - Submitted the report and ended the exam

  • August 12 1120 - Received the news that I passed

Personal Review

Hacking AI Or Hacking With AI?

The exam was super fun! However, I felt that the attacks against AI were fairly limited - I was expecting a bit more AI-hacking. Instead of hacking AI it felt more like using AI to hack. Don't get me wrong, this is still very useful, but based on the course materials my expectations were different.

Difficulty

I came into the exam severely underprepared. This was a failure from my side. I can't imagine passing other 300-level exams from OffSec (like the OSWE or OSEP) while going into them underprepared. On the other hand, OSWE and OSEP with AI wouldn't really feel 300-level either. A huge difference between this and other 300-level courses is the time - for this you only have 24 hours, whereas the others are 48 hours. But still, this is not the kind of exam where you tell Claude "Hey, hack this", you go for coffee and by the time you're back you passed. It is genuinely tough. If I were to do it again, I would set up a better AI harness which I would test more on several platforms and boxes. For anyone interested in this, check out Daniel Miessler's stuff - he is great! (see https://github.com/danielmiessler/fabric)

Cost

The course and exam bundle for the OSAI goes for the standard OffSec bundle price ($1749). The course materials are actually very good, so I would say that it is worth it. However, with AI changing every day, I am afraid that the course materials may quickly become outdated. This is something that can be verified 6-months from now.

There is also a hidden cost - the cost of AI. Most of us already have Claude/Codex/other AI subscriptions, however, for the exam you may need a higher tier. At least from my side, since I did not have an efficient harness set up in advance the costs were fairly high.

Conclusions

I highly recommend this course to anyone who wants to become more familiar with using AI offensively and with attacking AI systems. I wouldn't recommend this course for the LearnOne subscription (one-year access) since 90 days is more than enough to get ready for the course. Passing the exam definitely opens some doors - I received 500 LinkedIn connection requests the day I announced that I passed. Now keep in mind that this is a Red Team course - if you do not have hands-on experience with conducting attacks it may be quite tough to keep up with all the concepts, but if you have been doing Red Teaming or Penetration Testing for a while, you will not have any trouble. In case you're wondering what's next for me... Well, I will focus on catching up with some projects, continue delivering talks and doing research. Hopefully, in the next 6 months I will be able to start the OSED and achieve the OSCE3. Other than that, I will certainly apply the new skills I acquired from the OSAI. Happy hacking!

 
 
 

Comments


© HiveHack. All rights reserved.
Bee vigilant. Protect the hive.

Contact us to discuss your security needs.

Follow us on social media.

©2026 by HiveHack

  • Facebook
  • LinkedIn
  • YouTube
bottom of page