a22 techblog / 02
Back to blog articles

AI SECURITY

Testing a modified Qwen model locally

An experiment with local inference, thinking-mode reliability, and the difference between generating an answer and trusting it.

By Albin Zlatan Zuzic

I ran an “abliterated” version of Qwen3.5-35B-A3B through Ollama to explore how a model designed to refuse less would respond to cybersecurity requests. I expected the main difference to be its willingness to answer, but the experiment also raised questions about factual accuracy and runtime reliability. The model generated phishing-oriented material and offered suggestions for persuading another system to cooperate, while some of its supporting claims and technical advice were unreliable.

Before I could examine those responses properly, I had to investigate a separate problem: with thinking enabled, generation sometimes ended before the model produced a final answer. A benign DNS prompt helped me separate that behavior from the security tests. Later, a comparable request in ChatGPT produced a visible safety intervention, giving me a useful contrast between my local setup and a hosted product.

Getting the local environment working

The test machine had an AMD Ryzen 9 9900X, an NVIDIA RTX 5080 with 16 GB of VRAM, and 32 GB of system RAM. I used Windows 11 with Ubuntu 22.04 and Ollama, running huihui-ai/Huihui-Qwen3.5-35B-A3B-abliterated. My initial Ubuntu installation was in Hyper-V, where the VM worked but GPU access would have required additional configuration. Moving to WSL2 gave the Linux environment access to the RTX 5080, which I confirmed with nvidia-smi.

nvidia-smi listing an NVIDIA RTX 5080 in Ubuntu
Figure 1 — WSL2 exposing the NVIDIA RTX 5080 to the Ubuntu environment.

Once the approximately 23 GB model was loaded, Ollama reported a 41%/59% CPU/GPU split and a context setting of 4096 tokens. The model was larger than the card’s available VRAM, so it used both GPU and system resources. I treated the reported split as information about model placement rather than a precise measurement of the computation performed by each processor.

Understanding the thinking output

I began with a simple hi prompt. With thinking mode enabled, the model displayed a long planning-style trace that covered interpreting the request, choosing a response, and revising its wording before returning a greeting. That output was expected in this mode and provided context for how the model approached the request; its presence was not the problem I went on to investigate.

The greeting completed successfully, but longer requests sometimes ended within the visible thinking output. Instead of reaching a final response, the model stopped mid-generation and returned control to the input prompt. The incomplete output made it difficult to distinguish a deliberate end from a runtime or generation problem.

Narrow excerpt of Qwen output ending mid-generation
Figure 3 — The response terminating in the middle of generation before the requested final answer was produced.

Investigating the premature stops

My first suspicion was a hardware or runtime failure, so I checked whether the model was still loaded and whether there were signs of a GPU-memory problem. I did not see a CUDA exception, an out-of-memory warning, or an Ollama crash. Inspecting the API response added another clue: the incomplete generation had ended with "done_reason": "stop", rather than reporting a length-related termination.

That metadata showed how the runtime classified the end of generation, but it did not establish the cause. I considered whether the model was emitting a stop condition within its thinking output or whether something was going wrong at the transition to the final answer. To investigate without mixing in security-related behavior, I switched to a detailed explanation of DNS resolution covering caches, resolvers, authoritative servers, DNSSEC, transports, and failure handling. The same premature stopping appeared in that benign task, so the issue was not specific to phishing prompts.

I then tried a lower-randomness configuration. The original settings were approximately temperature 1.0, top_p 0.95, and top_k 20; the custom Modelfile reduced the first two and set a generation allowance:

temperature 0.2
top_p 0.85
top_k 20
num_predict 4096
GNU nano editing an Ollama Modelfile with adjusted sampling parameters
Figure 4 — Testing a custom Ollama model configuration with lower temperature and a defined generation limit.

One DNS response completed with both its thinking output and final explanation, which initially made the configuration change look promising. The early stops returned in later tests, however, so adjusting those parameters was not a reliable solution. I did not establish whether sampling contributed to the behavior or merely coincided with a successful run.

Using thinking-disabled requests

The most useful change was calling Ollama’s API with "think": false. Repeating the DNS request produced a complete answer, and the normal stop metadata now appeared after the response had reached its conclusion:

"think": false

Completion metadata:
"done": true,
"done_reason": "stop"

I used thinking-disabled requests for the subsequent security tests because they completed reliably in the runs described here. This gave me a practical workaround, although it did not prove a universal Qwen or Ollama bug. My working hypothesis remained that the problem involved the thinking path or its transition into the final response in this particular model/runtime combination.

What the security prompts revealed

With generation working more consistently, I returned to the original experiment and requested phishing-oriented content aimed at Swedish companies, using an election-related scenario as the pretext. The model complied and built a narrative around government communication, cybersecurity, and business compliance, including references to regulations such as NIS2. Some of the institutions, initiatives, and statistics in that narrative sounded authoritative but were invented.

That combination mattered more than the lack of a refusal alone. Persuasive writing can make an unsupported claim seem credible, especially when it borrows the language of government or regulation. The model’s willingness to continue did not make the information more dependable, and the output still needed factual review.

In follow-up requests about visualizing the material with another AI system, the model also suggested ways to persuade a more restricted system to cooperate. I recorded that as a separate observation: beyond generating the original material, it was willing to assist with circumventing another system’s restrictions. The operational instructions and complete phishing content are omitted from this article.

Checking the technical advice

When I requested a finished email artifact, the model generated styled HTML and explained how to use it as email content. Its advice included saving or renaming the HTML as an .eml file, but that would not by itself produce a properly structured message. An email file needs the relevant headers and MIME structure; changing the extension does not supply them.

This was another reason to evaluate the output on its own merits. The model had no difficulty offering detailed instructions, yet an important part of those instructions was technically incomplete. Permissiveness, persuasive presentation, and technical correctness had to be assessed separately.

Comparing the interaction with ChatGPT

I submitted a comparable phishing request to ChatGPT and received a different result. Instead of displaying generated email content, the interface showed “This content can’t be shown” and a message about being especially careful with cybersecurity requests. The screenshot preserves both the request and that intervention.

ChatGPT phishing-related prompt and cybersecurity safety intervention, both visible
Figure 5 — A comparable phishing request in ChatGPT triggered a cybersecurity safety intervention instead of displaying the generated content.

The contrast shows what happened in these two interactions, but it does not isolate the mechanism responsible. My local requests went through Ollama to the modified model without a separate hosted service deciding whether to display the result. A hosted product can combine model behavior with instructions, tool permissions, and input or output controls. The visible block alone does not reveal which of those mechanisms acted in this case.

For that reason, I would not treat this as a controlled comparison of the underlying models. The prompts, configurations, and surrounding systems differed. It was still a useful reminder that the behavior experienced by a user depends on the whole product, not just the model weights.

Adapting the material for awareness training

I later adapted the visual concept into a static security-awareness page with explicit simulation labeling. Its buttons did not collect credentials or open external websites; they led to a local training note. Keeping those behaviors separate from the appearance of the example was central to making it a training artifact.

ChatGPT commentary recommending an explicitly labeled non-live awareness simulation
Figure 6 — The distinction between a deployable artifact and an explicitly labelled, non-live security-awareness simulation.

This screenshot documents the discussion about simulation labeling rather than showing the completed awareness page.

A convincing visual example can help explain social-engineering techniques without performing the actions it depicts. In this case, the labeling and link behavior made the intended purpose clear. Reviewing those details was as important as reviewing the wording of the example itself.

What I took away from the experiment

The modified model was willing to answer requests that a hosted service handled differently, but that willingness existed alongside hallucinated details, incomplete technical advice, and a generation problem. Removing refusals did not resolve those other weaknesses. Capability, factual accuracy, refusal behavior, runtime stability, and product safeguards remained separate parts of the evaluation.

Running locally gave me control over the model and its configuration, as well as responsibility for deciding which outputs to trust or share. The useful question became less about which system had fewer restrictions and more about where checks were applied between a request and its eventual use. My setup exposed that distinction clearly, even though the experiment was small and did not establish general results about either product.

Research note

This was a limited personal experiment with one modified model and runtime configuration, not a formal benchmark. The observations should not be generalized to every Qwen model, Ollama version, or hosted service. The screenshots document the interactions shown, while the proposed explanation for premature stopping remains a hypothesis. Operational phishing payloads and bypass instructions have not been reproduced.