The shortest way is list(dict.fromkeys(items)). It removes duplicates and keeps the order in which each item first appeared.
Why set is not enough
items = ["b", "a", "b", "c", "a"] list(set(items)) # order is not guaranteed, e.g. ['c', 'a', 'b'] list(dict.fromkeys(items)) # ['b', 'a', 'c']
A set has no order guarantee, so converting a list to a set and back can shuffle it. A dictionary keeps insertion order (guaranteed since Python 3.7), and its keys are unique, so building a dictionary from the items drops repeats and keeps the first position of each. The values (all None here) are ignored.
An explicit loop
seen = set()
result = []
for x in items:
if x not in seen:
seen.add(x)
result.append(x)This does the same thing and is easier to extend, for example when you want to log duplicates or stop early. It is also what to write if the items are not hashable (see below).
Deduplicate records by a key
For dictionaries or rows, decide what "duplicate" means. Use a key function:
def unique_by(rows, key):
seen = set()
for r in rows:
k = key(r)
if k not in seen:
seen.add(k)
yield r
list(unique_by(orders, key=lambda r: r["order_id"]))It keeps the first record for each order_id. If you want the latest version, sort first, or keep a dictionary and overwrite: {r["order_id"]: r for r in rows}.values() keeps the last record per id, still in order of first appearance of the key.
Things to watch
- Items must be hashable for
setanddict.fromkeys. Lists and dicts are not. Convert them to tuples or use a key such as a tuple of the important fields. - Equal-looking values of different types (
1and1.0andTrue) are equal and hash the same, so only the first survives. - For large data in pandas,
df.drop_duplicates()does this, and in SQL it isDISTINCTorROW_NUMBER.