All articles
Share

Social Engineering has a Grammar: You Have to Learn to Read It

July 27, 2026
Attack Techniques
July 27, 2026
Jordan Schoenherr
Scientist
Title
SHARE
SHARE
SHARE

Key takeaways

1. Social engineering has a grammar: attacks can consist of a limited set of social cues, but their sequence shapes how a request is understood.

2. Identity, authority, rapport, and urgency can become more persuasive when they appear in the order expected within a legitimate interaction.

3. Attackers may use “script substitution” to replace secure procedures with familiar social or business scripts that make risky actions feel routine.

4. Detection must examine how an interaction unfolds against the expected procedure — not just individual words, cues, or tactics.

5. Watch Jordan present his full research on this topic at BSides Las Vegas, or schedule time to meet with Humanix during Black Hat to learn more about how conversational structure can become a security signal.

The ‘art’ of social engineering requires understanding and responding to a target

Like sentences, attacks have a grammar: they are more constrained. If you understand the grammar, you can defend against social engineering.

The success of vishing calls and phishing emails does not occur in a vacuum; it exploits basic features of human psychology. Reading security blogs, you’ll often find general references to urgency, authority, and rapport-building. These are Cialdini’s greatest hits, but their discussion is usually oversimplified. No single cue unlocks compliance.

‘Pretexts’ are the black boxes of cybersecurity. We can only see what goes inside them (words and intonation) and what comes out (compliance). Tactics, techniques, and procedures (TTPs) suggest there’s far more to them, but these taxonomies are inert lists rather than living attacks. We have to address the gap between the ‘art’ and ‘science’ of social engineering.

That's why it makes sense to look for something like a grammar of social engineering: sets of rules that are both structured and flexible.

How grammar and cues constrain meaning

A sentence isn't a bag of words. "Hacker attacks analyst” and "analyst attacks hacker" contain the same lexicon but have an opposite meaning. Words do carry a unique meaning, but the order is critical.

Figure 1. Attackers attempt to use cues to appear to follow a formal, normalized business process. Defenders must be aware of the appropriate procedures to defend against it.

Social engineering should work the same way. Social cues dot our sentences. “I’m from the Bank of X, we have noticed suspicious activity on your account” asserts authority within seconds. When legitimate, this statement establishes a person’s status and motivations - and the expectations of the call recipient. None of these pieces is dangerous alone; they need to work together. When distributed properly, "I'm calling from IT" justifies the request "can you read me the code on your screen".

It’s a mistake to assume our intuitions are always (or mostly) right. Deception cues are hard to find, but people are sensitive to them. If you put the right social cues in the wrong order it can alert someone’s mental SOC. The call recipient can have an uncanny experience. But this feeling can be triggered by an inconsistency, not necessarily the correct one.

Training is based around this idea: If we present the right cues and examples, then they can detect them. The problem is that training focuses on only a few cues, often ignoring the larger pattern. In the same way that we cannot learn French, Spanish, or Korean from a single class, we can’t expect a training session to build human lie detectors.

Recursive two-node cycle grammar Two boxes connected by a cycle. The arcs carry the verbs, while the boxes carry the noun phrases. An exit arc leaves the cycle for an accept state. requests information from replies to S0 the attacker the target breach start accept
Recursion depth 2 rounds


Figure 2.
Conversational exchanges can continue indefinitely. They require that the two parties take the other's perspective, requesting and providing information.

Structuring the Social

Humans are cognitive misers. If we can reduce our mental workload, we will. To make our lives easier, we create schemas in memory for common kinds of people and situations. CEOs have authority. Deals are scarce and time-dependent. We also make rules to complement these: if I’m speaking to a CEO, then I will comply; if a deadline is impending, I need to act on it.

This means social engineering would be defined by a small, closed set of social cues, schemas, and scripts the same way language works. These representations are the building blocks of a pretext. People establish their identity, make a request, and expect an answer. Throughout, they try to build rapport to create a sense of security and familiarity. When the time is right, urgency can be introduced and escalate the request until the target complies or they exit. Even if they don’t hang up, this passive compliance allows attackers to play the long game.

Each of these steps in a conversation looks normal. If we try to focus on deception and persuasion cues on their own, we waste our time: forced laughter, friendliness, and compliments are all part of professional environments. Supervisors try to persuade you to work on the weekend, salespeople say the shirt looks good on you, or your co-worker asks for help.

Why ordering is a social engineering signal

Persuasion literature has emphasized that order matters. It works because of basic memory principles: what comes first anchors interpretation of what follows, and what comes most recently is easiest to recall. That’s where cues, schemas, and scripts are stored. We should see that ordering influences social engineering.

In persuasion, the foot-in-the-door effect uses this sequential approach. Instead of making the big ask first, the social engineer starts small. It creates subtle mental inertia, causing drift toward the main target. Each response and concession keeps the person engaged. It works in the lab and in the wild. Reverse the order and compliance collapses.

When I analyzed hundreds of vishing call transcripts, I searched for sequences. Fixed scripts would be too obvious. Would-be victims would hear the mechanical, depersonalized speech and hang-up. This is why early robocalls failed. Savvy social engineers will adapt with persuasion principles, but framing and conversational logic still requires something akin to a grammar of social engineering. Adversaries don’t need to understand this explicitly; intuition can still succeed.

Figure 3. Secure procedures require establishing the identity of a caller before completing a requested action. Attackers can use social cues at key decision points to bypass verification.

Social engineers must make a request. Victims must respond. They can comply by providing information or performing a task or can stay on the line while the social engineer makes their case. They can also refuse or exit. But situations remain ambiguous. Even if the caller is questionable, people tend to avoid confrontation. Most resistance relies on going silent and disengaging. Confrontation is rare. Trail-and-error can help adversaries learn what works and what doesn’t. This only costs them time.

This is also why scripted, unsophisticated pretexting can still succeed. A help desk agent or CISO can know every tactic in the social engineering playbook and still comply because they arrived in the sequence a normal business call is supposed to arrive in. Combine this with personal and professional stressors and half the stage is set for any would-be social engineer.

How the trick works: Script substitution

Everyone knows about the bait and switch technique: a person offers something but provides something else. This only requires that people have short attention and memory spans. If the two offers look similar enough, they switch might go unnoticed.

To reduce their mental workloads, people rely on heuristics that substitute an easy problem for a hard problem. This is how certain tactics in social engineering work, a process I refer to as script substitution: an attacker replaces a security script with a social script. From service workers, their role is to help. For customers, they want help. In both cases, the quickest solution is to rely on a familiar script and let an authority guide us through the process.

When we read vishing scripts, we see that attackers adopt an identity congruent with what they want: legal, banking, customer service, and technical support. When this occurs first, the request that follows comes naturally. If they are a legal authority, they will ask questions consistent with that role. If suspicions do not make them stop and think, the caller is primed to think about the expected next steps. As time passes, the associated ideas – authority, urgency, etc. – surface. Even if a caller starts to question the interaction, most of what they find in their recent memories are consistent: they might not know what is required, and they might not remember what was said early in the conversations.

But conversations are exchanges. As long as the attacker can keep the caller on the phone, this passive compliance helps them build a case. If the attacker fails to execute a pretext convincingly, there might still be time to recover as long as they can keep the caller focused on their script. If customers don’t know what’s expected of them, an attacker can use a familiar script in place of an organizational procedure.

Why we must build for grammar, not just content

If language and social engineering are based on a small set of units, decreasing our certainty by constraining order, then deliberation will bypass these steps. Attackers have been exploiting that second fact for as long as they've been talking to people. Defenders have mostly been building tools that look at the units. It might be time to build ones that look at the sentence.

Turning the method into detection logic is not easy. Nor do all the lessons transfer. Like insider threat research, organization-based cases are hard to come by. Most believe that reputations will be adversely affected by disclosing breaches, let alone providing exact details of an attack. What data is available is not always structured effectively. Without appropriate taxonomies, we cannot obtain the needed telemetry.

This requires something far more complex than vishing training and more sophisticated than keyword analysis. Consider a few approaches:

Log the steps, not just the words

How the social cues are used is as important as what they are. Whatever transcript analysis or conversational AI a security team uses, they should examine call structure: which cues were present, in what position, and their relative frequency and location to others. A transcript full of the "right" words in the wrong order can provide an early warning. But AI-supplemented attacks are likely to erode this signal’s validity.

Separate language conventions from social engineering

Language and communications conventions have their own structure. When analyzing transcripts, be aware of word order. Languages like English follow a subject-verb-object structure and are relatively context-independent. Languages like Korean follow a subject-object-verb structure where speakers must attend to context. Korean also uses honorifics that imply status differences based on age and seniority.

Establish the Ordinary Grammar of Each Channel

Help desk calls, vendor payments, and recruitment are defined by their own sequences. Emails can be fragmentary and occur in multiple, parallel threads. Calls are more linear and contain emotional cues absent from emails. Organizational procedures and policies can – and should – dictate what occurs in each channel. Businesses such as those in banking provide emails to customers about these procedures, but they can easily be intercepted and integrated into pretexts.

How to learn more

Jordan will present the full version of this research, along with additional findings, at BSides Las Vegas on August 4th. You can also book a 1:1 meeting with him, as well as Humanix CEO Keith Stewart, onsite in Las Vegas during Black Hat from August 3rd-9th. More information is available here.

What is feature engineering

In practice, feature engineering is both science and a bit of witchcraft. It often involves both iteration and experimentation to uncover hidden patterns and relationships within the data. For instance, a data scientist might transform raw sales data into features such as average purchase value, purchase frequency, or customer lifetime value, which can significantly boost the performance of a churn prediction model. By thoughtfully engineering features, practitioners can provide machine learning models with the most informative inputs, ultimately leading to better accuracy and more robust predictions.

What’s more?

  • Incorporate more and more data sources
  • Feature engineering platform

What is data engineering

As we mentioned above, feature engineering is certainly a subset of data engineering. It involves the ingestion of data from a source, applying a series of transformations, and making the final result available to be queried by a model for training purposes. You can construct feature engineering pipelines to resemble data engineering pipelines, having schedules, specific source and sink destinations, and availability for querying. However, this configuration would only really apply once you have surpassed the experimentation stage and determined a need for a consistent flow of new feature data.

What is feature engineering

Image description

1. Functions

Functionally, there is nothing to differentiate data vs features - data points (link). Where feature engineering and data engineering really differ is in the objectives and motivations for constructing the pipelines. In general, data engineering serves a broader, more unified purpose than feature engineering. Data engineering platforms are constructed to be flexible and universal, ingesting various types and sources of data into a unified storage location where any number of transformations and use cases can be applied. The intent of a well constructed fact table or gold layer in a data lake is to provide a single source of truth that answers many different questions, produces many reports, and can be consumed by many downstream customers.

2. Practise

And in practice, an organization’s data engineering team will be responsible for the curation and maintenance of all data pipelines, not just those that relate to machine learning. These pipelines may power BI dashboards used by C-Suite, auditing reports that feed payroll, or event logs that show a user’s history of actions within the application.

Feature engineering, on the other hand, serves a specific purpose, finding the tailored inputs and columns that will generate the best predictive results for a machine learning model. Data scientists and machine learning engineers are not tasked with developing a universal data model that will ingest all data points throughout an organization, they just need to select, curate, and clean the data needed to power their models.

3. Machine learning

Now, as machine learning teams grow and begin to incorporate more and more data sources into their models, their feature engineering platform may start to resemble a larger data engineering platform in the tools and methodologies they employ. But, the intent is not to establish flexible data models that can be used throughout the organization - it is simply to power their machine learning models.

Enter your contact info and we'll be in touch soon

Oops! Something went wrong while submitting the form.