UK Phone Number Regex for JavaScript and Python (Tested)
Strip spaces, turn +44 into 0, then check for 07 plus nine digits. Tested JavaScript and Python code for every UK number type, E.164 storage, display spacing and libphonenumber.
The best way to remove duplicates from a list in Python is list(dict.fromkeys(items)). It keeps the first copy of each value and keeps the values in their original order, and it's one line with no imports. list(set(items)) also removes duplicates and is slightly faster, but a set has no order, so the result can come back shuffled. Use sorted(set(items)) if you want the unique values in alphabetical or numeric order instead. Both approaches need the items to be hashable, so for a list of dictionaries or a list of lists you have to build a hashable "fingerprint" for each item, or de-duplicate by one key such as an id.
In this guide I test every common version of the problem: plain lists of strings and numbers, removing duplicates while ignoring capital letters and stray spaces, lists of dictionaries (whole-dictionary duplicates, by one key, keeping the first or the last), lists of lists, tuples and objects, and pandas. I also timed each method on 100,000 items, so you can see how much the choice actually matters. Everything was run on Python 3.13.5 and pandas 3.0.6 on 6 October 2026, and the output under each example is pasted from the terminal.
Here's a list of UK towns with repeats, de-duplicated four ways:
# basics.py: remove duplicates from a list, with and without keeping order
towns = ["Leeds", "York", "Leeds", "Bath", "York", "Derby", "Bath"]
print("original: ", towns)
print("dict.fromkeys: ", list(dict.fromkeys(towns))) # keeps first-seen order
print("set (unordered): ", list(set(towns))) # order not guaranteed
print("sorted(set): ", sorted(set(towns))) # unique AND alphabetical
# The same thing as a loop, so you can see what dict.fromkeys does
seen = set()
unique = []
for town in towns:
if town not in seen:
seen.add(town)
unique.append(town)
print("seen-set loop: ", unique)
# How many duplicates were there?
print("removed", len(towns) - len(unique), "duplicates")
$ PYTHONHASHSEED=2 python3 basics.py
original: ['Leeds', 'York', 'Leeds', 'Bath', 'York', 'Derby', 'Bath']
dict.fromkeys: ['Leeds', 'York', 'Bath', 'Derby']
set (unordered): ['York', 'Derby', 'Bath', 'Leeds']
sorted(set): ['Bath', 'Derby', 'Leeds', 'York']
seen-set loop: ['Leeds', 'York', 'Bath', 'Derby']
removed 3 duplicates
What each line shows:
dict.fromkeys(towns) builds a dictionary whose keys are the towns. A dictionary can't have the same key twice, so later repeats are ignored, and since Python 3.7 dictionaries are guaranteed to remember insertion order. (The Python docs on dictionaries state this guarantee: "Dictionary order is guaranteed to be insertion order.") Wrapping it in list() gives back ['Leeds', 'York', 'Bath', 'Derby']: the order in which each town first appeared.set(towns) removes duplicates too, but here it gave ['York', 'Derby', 'Bath', 'Leeds'], which is neither the original order nor alphabetical. I fixed Python's hash seed (PYTHONHASHSEED=2) for this one run so the output in this post stays the same when I rebuild it; without that, the set's order can change every time you run the script, as the next section shows.sorted(set(towns)) is the right tool when you don't care about the original order and want the unique values sorted.dict.fromkeys does for you, written out by hand. It's longer, but it's the version you can bend to fit awkward cases, such as matching case-insensitively or by one field, which the later sections rely on.If you're new to the one-line style, a list comprehension can't do this job on its own, because it can't see what it has already added. The list comprehension examples post explains why, and when a plain loop is clearer.
A set stores items by their hash, not by position. For strings, Python adds a random "salt" to the hash each time it starts (hash randomisation, a security feature), so the order of a set of strings can change from one run to the next. To prove it, I ran the same line with three fixed seeds:
$ PYTHONHASHSEED=1 python3 -c 'print(list(set([...])))'
['Leeds', 'Bath', 'Derby', 'York']
$ PYTHONHASHSEED=2 python3 -c 'print(list(set([...])))'
['York', 'Derby', 'Bath', 'Leeds']
$ PYTHONHASHSEED=3 python3 -c 'print(list(set([...])))'
['Derby', 'York', 'Leeds', 'Bath']
$ PYTHONHASHSEED=1 python3 -c 'print(list(dict.fromkeys([...])))'
['Leeds', 'York', 'Bath', 'Derby']
Same input, three different orders. The last line, with dict.fromkeys, is the same every time. This is why code that uses list(set(…)) can pass on your machine and produce a differently ordered CSV, report or web page in production. If the order matters at all, don't use set() on its own. Small integers often appear to keep their order in a set, because an integer's hash is the integer itself. That's an implementation detail, so don't rely on it.
Real data is messier than towns typed by one person. Email addresses are a classic example: "Mahir@Example.com" and "mahir@example.com " (with a trailing space) are the same mailbox, but Python treats them as different strings. To match them, compare a cleaned-up key while keeping the original value:
# casefold.py: remove duplicates ignoring case and spaces, keep first spelling
emails = [
"Mahir@Example.com",
"mahir@example.com ",
"amelia@example.co.uk",
"AMELIA@EXAMPLE.CO.UK",
"oliver@example.org",
]
def unique_by(items, key):
"""Keep the first item for each key(item), in original order."""
seen = set()
out = []
for item in items:
k = key(item)
if k not in seen:
seen.add(k)
out.append(item)
return out
print("exact match only:", list(dict.fromkeys(emails)))
print("ignoring case: ", unique_by(emails, lambda e: e.strip().casefold()))
# casefold() vs lower(): casefold also folds the German sharp s
print("'Straße'.lower() =", "Straße".lower())
print("'Straße'.casefold() =", "Straße".casefold())
$ python3 casefold.py
exact match only: ['Mahir@Example.com', 'mahir@example.com ', 'amelia@example.co.uk', 'AMELIA@EXAMPLE.CO.UK', 'oliver@example.org']
ignoring case: ['Mahir@Example.com', 'amelia@example.co.uk', 'oliver@example.org']
'Straße'.lower() = straße
'Straße'.casefold() = strasse
dict.fromkeys on its own removed nothing, because no two strings were exactly identical. The unique_by() helper strips spaces and compares with casefold(), and it keeps the first spelling it saw for each address, so the output is three addresses rather than five.
Use casefold() rather than lower() when matching text. It's designed for case-insensitive comparison and handles more languages: as the output shows, "Straße".lower() keeps the German ß, while casefold() turns it into "strasse", so it matches "STRASSE". For English text the two usually behave the same, but casefold() is the correct choice and costs nothing.
unique_by() is worth keeping in a utilities file, because it takes any key function. key=str.casefold ignores case, key=lambda s: s.replace(" ", "") ignores spaces (handy for UK postcodes like "LS1 4AP" vs "LS14AP"), and key=lambda d: d["id"] works on dictionaries. If you're cleaning addresses before checking them, my post on email validation with regex covers the checking side; it's JavaScript, but the patterns carry over to Python's re module almost unchanged.
This is the most-searched version of the problem, and the first attempt nearly always fails with an error:
# dicts.py: remove duplicates from a list of dictionaries
orders = [
{"id": 101, "customer": "Ayesha", "total": 12.50},
{"id": 102, "customer": "Oliver", "total": 8.00},
{"id": 101, "customer": "Ayesha", "total": 12.50}, # exact duplicate
{"id": 103, "customer": "Ayesha", "total": 30.00},
{"id": 102, "customer": "Oliver", "total": 9.99}, # same id, new total
]
# 1. set() fails: dictionaries are not hashable
try:
set(orders)
except TypeError as e:
print("set(orders) ->", type(e).__name__ + ":", e)
# 2. Whole-dictionary duplicates: use a hashable fingerprint
def dedupe_dicts(items):
seen = set()
out = []
for d in items:
fingerprint = tuple(sorted(d.items()))
if fingerprint not in seen:
seen.add(fingerprint)
out.append(d)
return out
print("\nexact duplicates removed:")
for d in dedupe_dicts(orders):
print(" ", d)
# 3. Duplicates by one key, keeping the FIRST one seen
def dedupe_by_key_first(items, key):
seen = set()
return [d for d in items if not (d[key] in seen or seen.add(d[key]))]
print("\nby 'id', keep first:")
for d in dedupe_by_key_first(orders, "id"):
print(" ", d)
# 4. Duplicates by one key, keeping the LAST one seen (newest data wins)
def dedupe_by_key_last(items, key):
return list({d[key]: d for d in items}.values())
print("\nby 'id', keep last:")
for d in dedupe_by_key_last(orders, "id"):
print(" ", d)
# 5. Duplicates by two keys: one order per customer per total
print("\nby ('customer', 'total'):")
seen = set()
for d in orders:
k = (d["customer"], d["total"])
if k not in seen:
seen.add(k)
print(" ", d)
$ python3 dicts.py
set(orders) -> TypeError: unhashable type: 'dict'
exact duplicates removed:
{'id': 101, 'customer': 'Ayesha', 'total': 12.5}
{'id': 102, 'customer': 'Oliver', 'total': 8.0}
{'id': 103, 'customer': 'Ayesha', 'total': 30.0}
{'id': 102, 'customer': 'Oliver', 'total': 9.99}
by 'id', keep first:
{'id': 101, 'customer': 'Ayesha', 'total': 12.5}
{'id': 102, 'customer': 'Oliver', 'total': 8.0}
{'id': 103, 'customer': 'Ayesha', 'total': 30.0}
by 'id', keep last:
{'id': 101, 'customer': 'Ayesha', 'total': 12.5}
{'id': 102, 'customer': 'Oliver', 'total': 9.99}
{'id': 103, 'customer': 'Ayesha', 'total': 30.0}
by ('customer', 'total'):
{'id': 101, 'customer': 'Ayesha', 'total': 12.5}
{'id': 102, 'customer': 'Oliver', 'total': 8.0}
{'id': 103, 'customer': 'Ayesha', 'total': 30.0}
{'id': 102, 'customer': 'Oliver', 'total': 9.99}
Let's go through it.
TypeError: unhashable type: 'dict' means Python can't use a dictionary as a set member or dictionary key, because dictionaries can change. If a dictionary changed after being put in a set, its hash would no longer match where it was stored. The same goes for lists. The fix is to turn each dictionary into something hashable that represents it.
tuple(sorted(d.items())) turns {"id": 101, "customer": "Ayesha", "total": 12.5} into a tuple of key-value pairs, sorted so that two dictionaries with the same contents in a different key order still match. Tuples are hashable, so they can go in the seen set. The output shows the one exact duplicate (order 101 listed twice) removed, while order 102 appears twice because its two copies have different totals, so they aren't exact duplicates. This works as long as the values are hashable; if a value is itself a list or dictionary, use json.dumps(d, sort_keys=True) as the fingerprint instead.
Often you want "one row per id", even if the other fields differ. There are two sensible answers, and you should choose deliberately:
dedupe_by_key_first): the first order 102 wins, with a total of 8.00. Use this when the first record is the trustworthy one, such as the original sign-up.dedupe_by_key_last): {d[key]: d for d in items} overwrites each id with the latest dictionary, so order 102 ends up with a total of 9.99. Because the key was first inserted early, it keeps its original position while taking the newest value. Use this when later rows are updates, such as a CSV export where amendments are appended at the bottom.The keep-first one-liner uses a well-known trick: seen.add() returns None, which is falsy, so d[key] in seen or seen.add(d[key]) both checks and records the key in one expression. It's compact, but if your team finds it too clever, use the plain loop from the first section. They do exactly the same thing.
Make the key a tuple: (d["customer"], d["total"]). Tuples of hashable values are hashable, so the same seen-set pattern works. If your data comes from a CSV file, de-duplicating on a tuple of columns is a common cleaning step. You can read the file with the standard library, as I showed in reading a CSV file in Python without pandas, then run these helpers on the rows.
The same rule applies to everything: the items must be hashable, or you must make a hashable key from them.
# nested.py: lists of lists, tuples and objects
from dataclasses import dataclass
pairs = [[1, 2], [3, 4], [1, 2], [2, 1]]
try:
list(dict.fromkeys(pairs))
except TypeError as e:
print("dict.fromkeys(list of lists) ->", e)
# Convert each inner list to a tuple, de-duplicate, convert back
unique_pairs = [list(t) for t in dict.fromkeys(map(tuple, pairs))]
print("list of lists:", unique_pairs)
# Treat [1, 2] and [2, 1] as the same pair
unordered = [list(t) for t in dict.fromkeys(tuple(sorted(p)) for p in pairs)]
print("order inside pair ignored:", unordered)
# Tuples are hashable already
points = [(51.5, -0.12), (53.8, -1.55), (51.5, -0.12)]
print("list of tuples:", list(dict.fromkeys(points)))
# Objects: frozen dataclasses get __eq__ and __hash__ for free
@dataclass(frozen=True)
class Module:
code: str
title: str
modules = [Module("CS101", "Programming"), Module("MA102", "Maths"), Module("CS101", "Programming")]
print("frozen dataclasses:", list(dict.fromkeys(modules)))
# A plain class without __eq__/__hash__ compares by identity, so nothing is removed
class Plain:
def __init__(self, code): self.code = code
def __repr__(self): return f"Plain({self.code!r})"
plain = [Plain("CS101"), Plain("CS101")]
print("plain objects:", list(dict.fromkeys(plain)))
$ python3 nested.py
dict.fromkeys(list of lists) -> unhashable type: 'list'
list of lists: [[1, 2], [3, 4], [2, 1]]
order inside pair ignored: [[1, 2], [3, 4]]
list of tuples: [(51.5, -0.12), (53.8, -1.55)]
frozen dataclasses: [Module(code='CS101', title='Programming'), Module(code='MA102', title='Maths')]
plain objects: [Plain('CS101'), Plain('CS101')]
map(tuple, pairs), de-duplicate, then convert back. Note that [1, 2] and [2, 1] are different, because order inside a tuple matters.[2, 1] becomes (1, 2) and counts as a duplicate. This is handy for things like "who has messaged whom" or matching fixtures where home and away don't matter.dict.fromkeys works directly. Here, duplicate map coordinates are removed.@dataclass(frozen=True) gets __eq__ and __hash__ automatically, based on its fields, so two Module("CS101", "Programming") objects count as equal. A plain class, by default, compares by identity: two separate objects are never equal, however similar their attributes. The last line shows nothing removed. Either make it a frozen dataclass, define __eq__ and __hash__ yourself, or de-duplicate with unique_by(objects, key=lambda m: m.code).If your data is already in a pandas DataFrame, use pandas' own tools rather than converting to lists and back:
# pandas_dedupe.py: drop_duplicates on a DataFrame and on a list
import pandas as pd
towns = ["Leeds", "York", "Leeds", "Bath", "York", "Derby", "Bath"]
print("pd.unique: ", list(pd.unique(pd.Series(towns))))
df = pd.DataFrame([
{"id": 101, "customer": "Ayesha", "total": 12.50},
{"id": 102, "customer": "Oliver", "total": 8.00},
{"id": 101, "customer": "Ayesha", "total": 12.50},
{"id": 102, "customer": "Oliver", "total": 9.99},
])
print("\ndrop_duplicates():")
print(df.drop_duplicates().to_string(index=False))
print("\ndrop_duplicates(subset='id', keep='last'):")
print(df.drop_duplicates(subset="id", keep="last").to_string(index=False))
print("\npandas", pd.__version__)
$ python3 pandas_dedupe.py
pd.unique: ['Leeds', 'York', 'Bath', 'Derby']
drop_duplicates():
id customer total
101 Ayesha 12.50
102 Oliver 8.00
102 Oliver 9.99
drop_duplicates(subset='id', keep='last'):
id customer total
101 Ayesha 12.50
102 Oliver 9.99
pandas 3.0.6
pd.unique() keeps first-appearance order, like dict.fromkeys. df.drop_duplicates() removes rows that are identical across every column, and subset="id", keep="last" gives the "one row per id, newest wins" behaviour from earlier. There's also keep="first" (the default) and keep=False, which drops every row that has a duplicate. Installing pandas just to de-duplicate a list isn't worth it, though. The standard library handles lists perfectly well.
I timed each approach on 100,000 random integers with 20,000 unique values. Each figure is the best of three runs:
# speed.py: timing five ways to de-duplicate 100,000 items (20,000 unique)
import random, timeit, sys
random.seed(42)
data = [random.randrange(20_000) for _ in range(100_000)]
def fromkeys(x): return list(dict.fromkeys(x))
def via_set(x): return list(set(x))
def seen_loop(x):
seen, out = set(), []
for i in x:
if i not in seen:
seen.add(i); out.append(i)
return out
def in_list(x):
out = []
for i in x:
if i not in out: # scans the whole list every time
out.append(i)
return out
assert fromkeys(data) == seen_loop(data)
for name, fn, n in [("list(set(x))", via_set, 20), ("list(dict.fromkeys(x))", fromkeys, 20),
("seen-set loop", seen_loop, 20), ("'not in' list loop", in_list, 1)]:
t = min(timeit.repeat(lambda: fn(data), number=n, repeat=3)) / n
print(f"{name:24} {t * 1000:9.1f} ms")
print("Python", sys.version.split()[0])
$ python3 speed.py
list(set(x)) 1.9 ms
list(dict.fromkeys(x)) 2.8 ms
seen-set loop 3.9 ms
'not in' list loop 3900.7 ms
Python 3.13.5
Your numbers will differ by machine, but the shape won't:
list(set(x)) is fastest, but loses order.list(dict.fromkeys(x)) is only a little slower and keeps order. For almost every real program, this is the one to use.set(), which is still only a few milliseconds here. Use it when you need a custom key.if i not in out loop on a list took seconds rather than milliseconds, more than a thousand times slower. Checking whether a value is in a list means scanning the list from the start, every time, so the work grows with the square of the list size. Checking a set or dictionary is a hash lookup and takes about the same time however big it gets. This loop is the version most beginners write first, and it's fine for 20 items, but it's the reason a script that worked in testing can grind to a halt on a real data file.If you only remember one thing from this post, make it this short decision list. I go through it every time I clean a list, and it covers nearly every case I've come across in coursework and real projects:
list(dict.fromkeys(items)). This is the default choice.sorted(set(items)). Sorting puts them in a predictable order, so the scrambled order of the set no longer matters.list(set(items)). In practice this is rare, because the difference is tiny.unique_by() helper with a key function.tuple(sorted(d.items())) fingerprints.drop_duplicates(), with subset and keep set deliberately.Whichever you pick, print len() before and after. A quick "removed 3 duplicates" line in your script's output is the easiest way to spot a key function that's matching too much or too little.
# mistakes.py: removing items while looping skips elements
nums = [5, 5, 5, 3, 3, 3]
for n in nums:
if nums.count(n) > 1:
nums.remove(n)
print("remove() inside for loop:", nums) # expected [5, 3]
# Fix: build a new list instead of editing the one you are looping over
nums = [5, 5, 5, 3, 3, 3]
print("new list instead: ", list(dict.fromkeys(nums)))
# NaN is never equal to itself, so it is never treated as a duplicate
nan = float("nan")
vals = [1.0, float("nan"), float("nan"), 1.0]
print("two separate NaNs:", list(dict.fromkeys(vals)))
print("same NaN object twice:", list(dict.fromkeys([nan, nan])))
# 1, 1.0 and True are equal and hash the same, so they count as duplicates
print("1, 1.0, True:", list(dict.fromkeys([1, 1.0, True, 2])))
$ python3 mistakes.py
remove() inside for loop: [5, 3, 3]
new list instead: [5, 3]
two separate NaNs: [1.0, nan, nan]
same NaN object twice: [nan]
1, 1.0, True: [1, 2]
remove() an item, everything after it shifts left one place, and the loop skips the next item. The test should have given [5, 3] but returned [5, 3, 3]. With other inputs the same code happens to work, which makes it worse, because it passes a quick test. Build a new list instead.list(set(x)) when the order matters. As shown above, string order changes between runs. Use dict.fromkeys.not in on a list for big data. It's correct but painfully slow on large inputs. Track what you've seen in a set.1, 1.0 and True are equal. They compare as equal and have the same hash, so dict.fromkeys([1, 1.0, True, 2]) gives [1, 2]. This rarely matters, but it can surprise you with mixed data from JSON. If it matters, use (type(x), x) as the key.float("nan") is never equal to itself, so two separate NaN values are both kept. Only the same NaN object appearing twice is de-duplicated, because Python checks identity before equality. In pandas, drop_duplicates() does treat NaNs as equal, which is one reason to use pandas for numeric data with gaps.When you print the results for a report, f-string formatting makes it easy to line up columns or show counts like "removed 3 duplicates".
Use list(dict.fromkeys(my_list)). Dictionaries keep insertion order (guaranteed since Python 3.7), and repeated keys are ignored, so you get each value once, in the order it first appeared. For a custom rule, such as ignoring case, use a loop with a seen set.
set() is slightly faster. In my test on 100,000 items both took only a few milliseconds, with set() about a millisecond ahead of dict.fromkeys. But set() doesn't keep the order, and for strings the order can change between runs. The difference is so small that dict.fromkeys is the better default.
Dictionaries aren't hashable, so set() raises TypeError: unhashable type: 'dict'. For exact duplicates, track tuple(sorted(d.items())) in a set. For duplicates by one field, track d["id"] in a set to keep the first, or use list({d["id"]: d for d in items}.values()) to keep the last.
Loop over the list and keep a set of s.strip().casefold() values. Only append a string if its cleaned version isn't in the set yet. This keeps the first spelling you saw and drops later variants such as different capitals or trailing spaces.
Use collections.Counter(my_list). It counts each value, and [x for x, n in Counter(my_list).items() if n > 1] lists the values that appear more than once. len(my_list) - len(set(my_list)) tells you how many extra copies there are in total.
// note
This is a learning note from studying the web. It is one small topic, written so I can remember it. It is not a course and not a claim that I have finished the subject.
If a sentence is wrong, say so from the contact page and name this title. Drafts never appear here. Related notes, when they exist, are other published posts, and the same sample rule applies to each of them.
Strip spaces, turn +44 into 0, then check for 07 plus nine digits. Tested JavaScript and Python code for every UK number type, E.164 storage, display spacing and libphonenumber.
One-line ellipsis needs white-space: nowrap, overflow: hidden and text-overflow: ellipsis. Tested fixes for flexbox, grid and tables, plus 2- and 3-line truncation with line-clamp.
git branch -d name deletes the local branch, git push origin --delete name deletes the remote one. Tested errors, pruning stale branches, bulk clean-ups and how to undo a delete.