The deep web and the dark web sound like the same shady place, but they are very different things. The deep web is simply everything a search engine can't show you, and you use it every day; the dark web is a small, deliberately hidden network you can only reach with special software.
Three layers of the web
It helps to split the web into three parts, each one defined by a different question:
| Layer | What defines it | Examples |
|---|---|---|
| Surface web | Search engines index it | Blogs, news, public shops |
| Deep web | Not indexed | Email, banking, school portals |
| Dark web | Needs special software | Onion sites reached with Tor |
The first two are about search engines: is this page in the index or not? The third is about how you reach it: can a normal browser connect at all? That is why the dark web is a slice of the deep web. Mainstream search engines don't index onion sites, so they are unindexed too, but most unindexed pages have nothing to do with Tor.
What puts a page on the surface
A page is on the surface web when a search engine has found it, read it and added it to its index. Google describes three stages:
- Crawling: automated programs called crawlers follow links from page to page and download what they find.
- Indexing: the search engine analyses each page and stores what it is about.
- Serving: when you search, it picks the most relevant pages from the index.
Anything a crawler can't reach, or is told not to index, never makes it past step one or two. That leaves a huge amount of the web out of the index, for ordinary reasons.
Why most pages stay unindexed
- They sit behind a login. A crawler has no account, so it can't see your inbox, your bank statements or your company's dashboards.
- They are generated by a form. A flight search result or a library catalogue lookup only exists after someone types a query. Crawlers follow links; they don't fill in forms.
- Nothing links to them. A page no one links to is hard for a crawler to discover at all.
- The site asks not to be indexed. Site owners can tell search engines to keep pages out of results.
A worked example: keeping a page off the surface
Say a school runs a staff portal. The login page is public, but the timetable and pupil records behind it are not. Two different mechanisms keep pages out of search results.
The first is a login. When a crawler requests the records page, the server sees no session and redirects it to the login page, so the crawler never sees the records. This is real protection.
The second is a hint to search engines. A page can carry a noindex rule in its HTML:
<meta name="robots" content="noindex">And a site can ask crawlers not to fetch whole sections with a robots.txt file at its root:
User-agent: *
Disallow: /staff/Both keep pages out of search results, which puts them in the deep web. Neither one hides anything from a person. robots.txt is a public file that anyone can read, and a disallowed or noindex page still loads for anyone who has the URL. If the records page relied on these alone, it would be deep web and still wide open. Only the login actually protects it.
That is the everyday deep web: the parts of normal websites that need a password or a query, or that their owners chose to keep out of search. None of it needs special software.
What makes the dark web different
The dark web is not defined by search engines. It is made of sites on overlay networks: networks that run on top of the normal internet but use their own addressing and routing. The best known is Tor, and its sites are called onion services, with addresses that end in .onion.
You can't open one in an ordinary browser, for two reasons:
- The address isn't in DNS. A
.onionaddress isn't registered like a normal domain. It is derived from the service's public key, which is why it is a long run of random-looking characters (56 of them for current onion addresses). Only Tor knows how to find the service behind it. - The route is hidden on purpose. Tor sends traffic through several volunteer-run relays so that no single point can see both ends of a connection.
How Tor hides who you are
When you use Tor Browser, it picks three relays and builds a circuit through them: a guard relay, a middle relay and an exit relay. It wraps your request in three layers of encryption, one per relay. Each relay can remove only its own layer, which tells it the next hop and nothing more.
Here is one request travelling through a circuit to an ordinary website:
A request travelling through a Tor circuit
Step 1 of 8: Tor Browser wraps your request in three layers of encryption, one for each relay.
Your internet provider sees that you are using Tor, but not which sites you visit. The website sees a Tor exit relay, but not who you are. That split is the whole point.
Onion services hide the server too
Visiting a normal website through Tor hides you, but the website's own address is public. An onion service hides both sides. Instead of using an exit relay, the visitor picks a relay to act as a rendezvous point, and the visitor and the service each build their own circuit to it and meet there. The connection stays inside the Tor network from end to end, so neither side learns the other's IP address.
This is what the video means by sites that 'hide identities more carefully'. It is also why onion services are slow compared with the normal web: every request crosses several relays run by volunteers around the world.
Who uses the dark web, and why
The dark web has a reputation, and part of it is earned. Its anonymity attracts scams, markets for stolen data and illegal trade, and those sites are hard to shut down because their servers are hidden.
The same anonymity protects people with good reasons to hide:
- Privacy: people who don't want advertisers, their internet provider or anyone watching the network to build a record of what they read.
- Getting around censorship: people in countries that block news sites or social networks can reach them through Tor.
- Whistleblowers and journalists: some news organisations run onion services so that sources can contact them without revealing who they are.
The technology is neutral. What matters is what a particular site does with it.
Common mistakes
- Treating 'deep web' and 'dark web' as the same thing. Your email is deep web. It isn't dark web, and it never touches Tor.
- Thinking 'not in Google' means 'private'. An unlinked page, a
noindextag or arobots.txtrule only keeps a page out of search results. Anyone with the URL can still open it. Protect sensitive pages with authentication. - Assuming Tor makes you invisible. Tor hides your IP address, not your behaviour. Log in to an account over Tor and that site knows exactly who you are. On an ordinary website without HTTPS, the exit relay can read your traffic.
- Believing everything on the dark web is illegal. Much of it is, but onion services are also used by privacy tools, news organisations and people avoiding censorship.
Key takeaways
- The deep web is anything search engines don't index: logins, form results and pages kept out of search. Most of it is everyday, private stuff.
- The dark web is a small part of the deep web, on overlay networks like Tor, that you can't reach with a normal browser.
- Tor sends traffic through three relays with layered encryption, so no single relay knows both who you are and where you are going.
- Keeping a page out of search is not security: only authentication protects it.
- Not all deep web is dark, and not all dark web is criminal.