Skip to content

Alex Carr

Data & Solutions Architecture from the field.

  • Home
  • Blog
  • About

Three People Are Spider-Man. Your System Thinks They’re the Same Person.

Alex CarrOctober 1, 2026October 1, 2026 No Comments
Three different people in red caps walk through separate turnstiles into a building lobby. An access-control panel beside them shows all three matched to the identity "Spider-Man" with confidence scores of 92%, 89% and 87%, and a red alert reading "Identity conflict: 3 people matching the same identity".

How Spider-Man can teach us about identity, master data, and what data architects actually do.

Fair warning: this one is a walk-through rather than a quick read. Grab some popcorn and settle in.

Three Spider-Men walk up to Avengers Tower.

One is Peter Parker. Another is also Peter Parker from a different universe. The third is Miles Morales.

All three are legitimately Spider-Man.

The security system has a record for “Spider-Man”

Whoops! We may have a problem.

If the system treats Spider-Man as the identity, how does it know which person is standing at the door? Does every Spider-Man get the same access? If one Peter Parker loses access, do the others lose it too?

Or worse: what if the system thinks three different people are one person?

The problem isn’t that three people can be Spider-Man. The problem is that our system thinks Spider-Man is a person.

Welcome to identity resolution.

This will help illustrate the topic and how data architects help resolve an issue like the above.

Why this matters

At first glance, this may sound like a data quality problem and some may just say “Pick one Spider-Man, clean the duplicates and move on”

Biggest problem is these might not be duplicates and we definitely don’t want to get rid of a real Spider-Man

All three people may legitimately exist. They may even legitimately hold the same role. What matters is whether our systems understand that they are different people.

If the system thinks three people are one, the wrong Peter Parker might inherit another Peter’s access, approvals, account information, or history.

A split panel. Left: "False positive — 3 people to 1 identity, wrong person gets access", showing three people matched to a single identity and a disguised figure granted entry. Right: "False negative — 1 person to 3 identities, access survives termination", showing one person whose HR record reads terminated while physical access and IT records remain active.

That’s a false positive: we matched records that shouldn’t have been matched.

The opposite can be just as dangerous.

Suppose our Peter Parker exists three times — once in HR, once in the identity system, and once in physical security.

HR terminates Peter and his HR record becomes inactive but the badge system thinks it’s Peter Parker is someone else.

Now we have a badge that stayed active and access still existing that should have been removed.

That’s a false negative: records belonging to the same person weren’t connected.

At this point the duplicate data isn’t just making a report look ugly it has become a security problem.

This is where Master Data Management (MDM), identity management, access control, and architecture begin to intersect.

MDM itself shouldn’t decide whether Spider-Man can enter Avengers Tower.

It should make sure the system making that decision knows exactly which Spider-Man is standing outside.

So how do we build that? Let the fun begin.

Step 1: Figure out who knows what

Before designing anything, we need to understand the systems that already exist.

Maybe Avengers Tower has an HR system, an identity directory, a physical security system, suit inventory, training records, and a heroic incident-management platform that I’m sure Tony Stark dramatically overengineered.

If we’re lucky, all of that is documented and current but we’re usually not that lucky.

So we find the people responsible for those systems and start asking questions.

What was this system built to answer?

HR was built to answer employment questions.

A badge system was built to answer questions about physical access.

A directory manages digital identities and accounts.

None of those systems is necessarily wrong, they’re simply designed to answer different questions and that distinction matters.

How does a change get into each system?

When something changes, what happens?

Does someone update it manually? Does it receive data from another application?

Does that happen immediately, hourly, nightly, or whenever someone remembers to upload a spreadsheet?

Timing can matter. Imagine an access-control system that receives termination information once every 24 hours.

The data may be correct but that is a long time to still allow access.

What do users already not trust? Ask directly

System owners usually know exactly which parts of their data are unreliable.

Maybe addresses are rarely updated or contractor end dates aren’t consistently maintained.

Maybe everyone knows not to trust the phone-number field because nobody can remember where it comes from.

This isn’t about assigning blame, it’s about understanding reality.

What we produce: a system inventory documenting what each system does, who owns it, how often it updates, known pain points, and what information it is genuinely authoritative for.

That’s the beginning of the architecture and we haven’t touched the data yet.

Step 2: Decide what a “person” actually is

This may sound like an easy or hard question depending on your thought process but either way it is one of the most practical decisions in the project.

Who belongs in our person master? Employees? Contractors? Applicants who were never hired? Visitors? Former employees? Customers?

What about an AI agent with its own credentials and access to company systems?

Let’s head back to Spider-Man.

Peter Parker is a person. Spider-Man is a role that Peter Parker holds.

Those are not the same thing.

Peter Parker on the left and Miles Morales on the right, connected to a shared badge labelled "Role: Masked Hero". Peter's side lists roles held over time — visitor, intern, employee, masked hero — and Miles's lists student and masked hero. A not-equal symbol sits between the two people, above the caption "Two people can hold the same role. They are still two people."

Miles Morales can also hold the role Spider-Man without becoming Peter Parker.

One Peter Parker may have multiple relationships with the organization over time.

Maybe he first visited Avengers Tower as a guest but later became an intern. Eventually he became an Avenger.

Those relationships changed, not the person.

That distinction becomes extremely important when we start connecting systems.

It can also matter from a security perspective. If someone leaves an organization under circumstances that restrict future access, we probably don’t want them returning six months later as a contractor simply because the contractor system created a brand-new identity.

Of course, what information an organization can retain, and for how long, must also account for applicable privacy laws, regulations, policies, and legitimate business purposes.

What we produce: a written scope defining what the organization considers a person, which populations are included, and how roles and relationships are represented.

Step 3: Decide who gets to be right

Now we’re getting into what most people picture when they hear “master data.”

Suppose we have identified Peter Parker correctly. Great but we still have a problem.

Several systems have information about him, and they don’t necessarily agree.

HR says:

Legal name: Peter Parker Employment status: Active Photo: Peter from three years ago

The identity directory says:

Account: PParker Email: peter.parker@avengers.example

Physical security says:

Badge: 61601 Building access: Avengers Tower, Levels 1–7

Suit inventory says:

Current suit: Red Suit Last checkout: Tuesday

Which system is the source of truth? Careful of this trap when putting together a solution.

There doesn’t have to be one system that is right about everything.

HR may be authoritative for Peter’s legal name and employment status.

The directory may be authoritative for his network account.

Physical security may be authoritative for his badge and access history.

Suit inventory probably knows more about what Peter is currently wearing than HR ever will.

Each system should be trusted for the information it is responsible for maintaining.

Four source systems feed a single glowing golden record for Peter Parker. HR supplies legal name, the directory supplies account ID, the security system supplies badge status, and suit inventory supplies current suit — each contributing only the attribute it is authoritative for.

In MDM, part of this becomes survivorship: once we’ve determined that multiple records represent the same entity, which values should survive into the mastered record?

That doesn’t always mean “take HR’s answer”.

It might mean: Use HR for legal name, the directory for account ID, physical security for badge status.

Depending on your systems it may even be right to take the most recent verified address regardless of system, just be careful not to fall into a loop.

The main point is it is not one size fits all and perfectly acceptable for your master data to come from multiple sources as long as they are authoritative of that data and if the primary source has no value, use an agreed fallback.

Now our master record can represent the best available version of Peter without pretending one application knows everything about him. That said, companies do build a centralized source of truth for downstream systems which may get its source of truth from various source systems.

What we produce: a source-to-attribute matrix showing each attribute, its authoritative source, update frequency, and fallback rules.

This may look like a boring spreadsheet but it is also one of the most valuable artifacts in the entire architecture.

Developers will thank you for it. Six months later, when someone asks why a field contains what it does, you’ll thank yourself for it too.

Step 4: Decide how we know two records are the same person

Now things get energetic.

Three records arrive:

Peter Parker Queens, New York Government ID: 12345

Peter Parker Queens, New York Government ID: 12345

Peter Parker Earth-96283 Government ID: 12345

Same name. Same identifier. Potentially different people.

Welcome to the multiverse.

Three candidate records for Peter Parker from different universes sharing a name and government ID but differing on address and universe. A matching-engine panel shows name and government ID matching while address and universe conflict, producing 74% confidence and a "human review required" flag, with a steward at a screen choosing merge, reject or review.

This is where the matching rules matter, and where we need to lean heavily on the business to understand what combinations of information actually establish identity.

Maybe the government ID is normally enough. Unfortunately, nobody told the DMV to account for the multiverse.

Maybe name + date of birth + address creates a probable match.

Maybe certain combinations should never automatically merge.

Then there’s everything the system genuinely cannot determine confidently. Those records need human review.

How wide we make that gap is a design decision.

Make the matching rules too narrow and we’re automatically merging people nobody examined.

Make them too wide and we’ve created a review queue so large that nobody will ever work it. That’s basically auto-rejecting with extra steps.

This is why the business has to participate in matching design.

Architects can design sophisticated algorithms but the people running the business understand how identities are actually created, changed, duplicated, and corrected.

And if they don’t trust the matching rules, they won’t trust the master.

What we produce: documented matching rules, thresholds for automatic matching and human review, exception handling, and a named owner for the review queue.

Because a review queue without an owner isn’t a process, it’s a backlog.

Step 5: Design for being wrong

You will be wrong sometimes and that is ok. Design for wrong while being wrong is still cheap.

Can you unmerge?

Suppose the system merged two Peter Parkers and six months later we discover they’re different people.

Can we separate them? Can we reconstruct what each record looked like before the merge?

If not, a matching mistake just became a permanent data problem.

Let’s go back to school and make sure you show your work.

Look at Peter’s mastered record.

Why does it say his address is in Queens? Where did that value come from? When did we receive it?

What rule caused it to win over another address?

This is field-level lineage, and it becomes extremely important when someone challenges the data or an auditor asks how a decision was made. Again, it’s ok to be wrong but make sure you can show systematically what happened because you definitely don’t want people thinking you are manipulating data on your own.

Who needs to know when something changes?

Suppose Peter and another Peter were incorrectly merged.

We fix the master record, excellent job! But five downstream applications already consumed the bad version.

Fixing the master without notifying or correcting downstream consumers doesn’t fix the problem.

What if you fixed a termination but the downstream access system was never informed. You just put your access control friends in the middle of something they never had control of and they are the ones getting blamed.

What we produce: an unmerge process, field-level lineage, auditability, and a defined method for notifying downstream consumers when mastered data changes.

A trustworthy master isn’t one that never makes mistakes. It’s one that can explain and recover from them.

Step 6: Design for the field nobody has asked for yet

Most people that have worked on data projects know that six months from now you are going to get the question “Can we add one more field?”

The business always does and that’s ok.

Maybe SHIELD suddenly wants to track Multiverse of origin.

Fair enough but before adding MultiverseOrigin to a table, don’t be afraid to ask questions.

Who is allowed to request the attribute? What does it actually mean?

Which system owns it? Who maintains it?

How often can it change? What happens when it’s unknown?

Which downstream systems need it? Who approves its use?

An attribute with no named owner and no update process may still be data but it can quickly become untrusted data.

It’s a column that will be accurate for about a quarter.

So build the process before the request arrives.

What we produce: a lightweight intake and governance process for adding or changing mastered attributes.

It doesn’t need to be a 40-page governance document. It just needs to answer the questions everyone otherwise waits until production to ask.

So what did we actually build?

Let’s look at what came out of those six steps:

  • A system inventory and ownership model
  • A definition of what counts as a person
  • A source-to-attribute and survivorship matrix
  • Matching rules and thresholds
  • A human-review process with an owner
  • An unmerge and recovery process
  • Field-level lineage
  • A downstream notification strategy
  • A process for introducing new attributes

Notice what’s mostly missing from that list.

Technology.

We haven’t selected a database. We haven’t argued about APIs. We haven’t chosen an MDM platform.

We haven’t drawn a cloud diagram with 47 boxes and arrows.

Yet we’ve made most of the decisions that determine whether the system will actually work.

That’s architecture.

This absolutley, does not mean the technology doesn’t matter and that diagrams should not be built. Visuals are always good and I am a strong advocate for them.

Just remember technology implements these decisions.

Give a development team an expensive MDM platform without answering these questions and they can build a technically beautiful system that confidently answers the wrong question.

Back to Avengers Tower

Three Spider-Men are standing outside.

Our architecture no longer cares that they share a title.

It knows that Peter Parker is a person. It knows Miles Morales is a different person.

It knows Spider-Man is a role either of them may hold.

It knows which systems are trusted for which attributes.

It knows how records are matched. It knows when it isn’t confident enough to make that match automatically.

It knows who resolves the ambiguity and if it makes a mistake, it knows how to unwind it.

The access-control system can now make its decision using the correct identity.

Most of what made that possible wasn’t code.

It was conversations. Definitions. Ownership.

Rules and decisions written down before someone started building.

The problem was never deciding which one was really Spider-Man.

The system just needed to know which one was standing at the door.

Post navigation

Previous: A Project List Is Not a Priority List

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Proudly powered by WordPress | Theme: Smart Portfolio by Code Work Web.