Information Gathering — Web Reconnaissance, DNS, Subdomains and Fingerprinting (HTB CWES)
Built from the Information Gathering – Web Edition module of the HTB Academy Web Penetration Tester path (HTB CWES). Revision notes: every tool and command kept, the concepts condensed, a full cheatsheet at the end. Labs and lessons live on HTB Academy.
Reconnaissance is the foundation of a web assessment — map the target before touching it. The goal is to uncover assets (pages, subdomains, IPs, technologies), hidden information (backups, config files, comments), the attack surface, and intelligence (people, emails) you can turn into an entry point.
1. Active vs passive recon
| Active | Passive | |
|---|---|---|
| How | Directly interact with the target | Use only public, third-party sources |
| Examples | Port scan, vuln scan, banner grab, vhost/subdomain brute force, crawling | WHOIS, DNS lookups, CT logs, search engines, Wayback, social media, GitHub |
| Detection | Logged, can trip IDS/WAF | Very low — looks like normal internet use |
Start passive to stay quiet, then go active where you need depth. Always have authorization before active recon.
2. WHOIS
WHOIS is the internet's phonebook: who registered a domain (or owns an IP block / ASN). Install and query:
sudo apt install whois -y
whois inlanefreight.com
A record yields the registrar, registrant/admin/tech contacts, creation/expiry dates, and name servers. For a pentester that's social-engineering fodder (names, emails, phones), infrastructure clues (name servers → hosting), and — via history services like WhoisFreaks — ownership changes over time. Red flags on a suspicious domain: very recent registration, privacy-masked registrant, bulletproof-host name servers.
3. DNS fundamentals
DNS turns names into IPs through a recursive chain of servers.
A zone is a slice of the namespace an authoritative server manages; its zone file holds the records. The hosts file (/etc/hosts on Linux/macOS, C:\Windows\System32\drivers\etc\hosts on Windows) overrides DNS locally — essential for reaching lab vhosts that have no public record:
<IP Address> <Hostname> [<Alias> ...]
10.129.42.190 inlanefreight.htb www.inlanefreight.htb
Record types to know:
| Type | Meaning |
|---|---|
A / AAAA |
Hostname → IPv4 / IPv6 |
CNAME |
Alias → another hostname |
MX |
Mail server(s) for the domain |
NS |
Authoritative name server(s) |
TXT |
Arbitrary text (SPF, verification, _1password=… leaks) |
SOA |
Zone admin info (primary NS, serial, timers) |
SRV |
Host + port for a service |
PTR |
Reverse lookup, IP → hostname |
Why it matters: records reveal subdomains, mail servers and infrastructure; a CNAME to a dead host hints at subdomain takeover; a new subdomain appearing over time is a new entry point; TXT records leak which SaaS the org uses.
4. Digging DNS with dig
dig (Domain Information Groper) is the workhorse. Companions: nslookup, host (simpler); dnsenum, fierce, dnsrecon, theHarvester (automated).
dig domain.com # default A lookup
dig domain.com MX # mail servers
dig domain.com NS # name servers
dig domain.com TXT # TXT records
dig domain.com SOA # zone admin record
dig @1.1.1.1 domain.com # query a specific resolver
dig +trace domain.com # full resolution path (root → authoritative)
dig -x 192.168.1.1 # reverse lookup (PTR)
dig +short domain.com # just the answer
dig +noall +answer domain.com # only the answer section
Read the output in four parts: header (status NOERROR, flags), question, answer (name TTL IN A IP), footer (query time, server). ANY is mostly ignored now (RFC 8482). Mind rate limits — excessive queries can be blocked.
5. Subdomain enumeration
Subdomains hide dev/staging boxes, admin portals, legacy apps and leaked data. Two approaches:
Active — zone transfer (AXFR). A misconfigured name server will hand you the entire zone (every subdomain + IP). Rare today, always worth a try:
dig axfr @nsztm1.digi.ninja zonetransfer.me
Active — brute force. Test a wordlist of names against the domain. Tools: dnsenum, fierce, dnsrecon, amass, assetfinder, puredns, ffuf, gobuster.
dnsenum --enum inlanefreight.com \
-f /usr/share/seclists/Discovery/DNS/subdomains-top1million-20000.txt -r
(-f wordlist, -r recursive.)
Passive — Certificate Transparency logs. Every SSL/TLS cert a CA issues is logged publicly; the SAN field lists subdomains, including dead/old ones. No wordlist, no guessing. Query crt.sh from the CLI:
curl -s "https://crt.sh/?q=facebook.com&output=json" \
| jq -r '.[] | select(.name_value | contains("dev")) | .name_value' | sort -u
(Censys is the heavier, filterable alternative.)
6. Virtual hosts
One IP can serve many sites; the web server picks which by the HTTP Host header. Non-public vhosts have no DNS record but still respond if you send the right Host — so vhost fuzzing finds them. gobuster (or ffuf/feroxbuster):
gobuster vhost -u http://inlanefreight.htb:81 \
-w /usr/share/seclists/Discovery/DNS/subdomains-top1million-110000.txt --append-domain
--append-domain is required in current Gobuster; add -t (threads), -k (ignore TLS errors), -o (save output). Note the difference from subdomains: vhost discovery works by varying the Host header against a known IP, not by DNS.
7. Fingerprinting
Identify the stack (server, OS, CMS, framework, WAF) so you can aim exploits precisely.
- Banner grabbing / headers — the
ServerandX-Powered-Byheaders, redirects, and paths like/wp-json/(WordPress):bash curl -I inlanefreight.com # HTTP headers only curl -I https://www.inlanefreight.com - WAF detection — before probing hard, know what's filtering you:
bash wafw00f inlanefreight.com # e.g. "behind Wordfence (Defiant) WAF" - Nikto — server scanner; fingerprint-only tuning:
bash nikto -h inlanefreight.com -Tuning b - Others:
whatweb,nmap(service/OS,-sV/-O, NSE scripts), and the GUI/online Wappalyzer, BuiltWith, Netcraft.
8. Crawling and hidden paths
Crawling (spidering) follows links from a seed page to map the site and harvest links, comments, metadata and sensitive files (.bak, .old, web.config, settings.php, logs, keys). Value comes from context — a /files/ directory in the links plus a "file server" comment is a lead.
robots.txt (site root) lists paths the owner wants crawlers to skip — which is exactly where the interesting stuff often is:
User-agent: *
Disallow: /admin/
Disallow: /private/
Allow: /public/
Sitemap: https://example.com/sitemap.xml
Read Disallow entries as a map of hidden directories (admin panels, backups).
.well-known/ URIs (RFC 8615) are standardized metadata endpoints — an IANA registry lists them. security.txt gives a security contact; openid-configuration dumps a JSON of OAuth/OIDC endpoints (authorization_endpoint, token_endpoint, jwks_uri, scopes) — a ready map of the auth surface:
curl -s https://example.com/.well-known/openid-configuration | jq
Automated crawlers: Burp Suite Spider, OWASP ZAP, Scrapy (Python framework), Apache Nutch. The module ships a custom Scrapy spider, ReconSpider, that outputs results.json (emails, links, external/JS files, images, comments):
pip3 install scrapy
python3 ReconSpider.py http://inlanefreight.com # → results.json
9. Search engine discovery (Google dorking)
Search operators turn Google into a recon tool (the Google Hacking Database has thousands of ready dorks):
| Operator | Finds |
|---|---|
site: |
Pages on one domain |
inurl: / allinurl: |
Term(s) in the URL |
filetype: / ext: |
Files of a type (pdf, sql, conf) |
intitle: / intext: |
Term in the title / body |
cache: |
Google's cached copy |
"…" · - · OR · * · .. |
Exact phrase · exclude · either · wildcard · number range |
site:example.com inurl:login # login pages
site:example.com filetype:pdf # exposed documents
site:example.com inurl:config.php # config files
site:example.com filetype:sql # database backups
site:example.com (inurl:admin OR inurl:login)
10. Web archives — the Wayback Machine
web.archive.org stores historical snapshots of sites since 1996. Passive and powerful: recover old pages, directories, files and subdomains that are gone from the live site (but maybe still reachable), diff how the site changed, and gather OSINT — all without touching the target.
11. Automating recon
Frameworks chain these tasks for speed, scale and consistency:
- FinalRecon — headers, WHOIS, SSL, crawl, DNS, subdomains, directories, Wayback, port scan.
- Recon-ng — modular framework (DNS, subdomains, ports, crawling, exploits).
- theHarvester — emails, subdomains, hosts, names, ports from public sources.
- SpiderFoot — broad OSINT automation across many data sources.
- OSINT Framework — a directory of sources/tools.
git clone https://github.com/thewhiteh4t/FinalRecon.git && cd FinalRecon
pip3 install -r requirements.txt && chmod +x ./finalrecon.py
./finalrecon.py --headers --whois --url http://inlanefreight.com
./finalrecon.py --full --url http://inlanefreight.com # everything
12. What to carry into the CWES exam
- Passive first, then active. WHOIS, CT logs (crt.sh), Google dorks and Wayback cost nothing and tip nobody off; brute force and crawling come after.
- Enumerate subdomains and vhosts — they're different: subdomains via DNS/CT logs, vhosts via the
Hostheader against the IP. Add every host you find to/etc/hosts. - Fingerprint before you attack —
curl -I,wafw00f,whatweb/nikto. Knowing it's WordPress behind Wordfence changes everything you do next. - Mine robots.txt, .well-known and the Wayback Machine for hidden paths and the auth surface.
- The skills assessment = whois + robots.txt analysis + subdomain brute force + crawl, with discovered subdomains added to your hosts file. Practice the chain end to end.
Cheatsheet — Information Gathering (Web)
WHOIS
sudo apt install whois -y
whois inlanefreight.com
DNS / dig
dig domain.com # A record
dig domain.com MX|NS|TXT|SOA|CNAME|AAAA
dig @1.1.1.1 domain.com # specific resolver
dig +short domain.com # answer only
dig +noall +answer domain.com # answer section only
dig +trace domain.com # full resolution path
dig -x 1.2.3.4 # reverse (PTR)
dig axfr @ns.target.tld target.tld # zone transfer (AXFR)
nslookup domain.com # host domain.com (simpler tools)
Subdomains
dnsenum --enum DOMAIN -f /usr/share/seclists/Discovery/DNS/subdomains-top1million-20000.txt -r
# CT logs via crt.sh:
curl -s "https://crt.sh/?q=DOMAIN&output=json" | jq -r '.[].name_value' | sort -u
# others: fierce, dnsrecon, amass, assetfinder, puredns, ffuf, gobuster dns
Virtual hosts
gobuster vhost -u http://TARGET_IP:PORT -w WORDLIST --append-domain [-t 50] [-k] [-o out.txt]
Fingerprinting
curl -I https://TARGET # headers / Server banner
wafw00f TARGET # detect WAF
nikto -h TARGET -Tuning b # fingerprint scan
whatweb TARGET # tech fingerprint
nmap -sV -O TARGET # service/OS
# GUI/online: Wappalyzer, BuiltWith, Netcraft
Crawling / hidden paths
curl -s https://TARGET/robots.txt
curl -s https://TARGET/.well-known/security.txt
curl -s https://TARGET/.well-known/openid-configuration | jq
pip3 install scrapy && python3 ReconSpider.py http://TARGET # → results.json
# crawlers: Burp Spider, OWASP ZAP, Scrapy, Apache Nutch
Google dorks
site:TARGET inurl:login
site:TARGET filetype:pdf (also: filetype:sql | ext:conf | ext:cnf)
site:TARGET inurl:config.php
intitle:"index of" | intext:"password" | cache:TARGET
# Google Hacking Database: exploit-db.com/google-hacking-database
Archives & automation
# Wayback Machine: https://web.archive.org/
git clone https://github.com/thewhiteh4t/FinalRecon.git && cd FinalRecon
pip3 install -r requirements.txt && chmod +x ./finalrecon.py
./finalrecon.py --headers --whois --url http://TARGET
./finalrecon.py --full --url http://TARGET
# frameworks: Recon-ng, theHarvester, SpiderFoot, OSINT Framework
/etc/hosts (add discovered hosts)
<TARGET_IP> inlanefreight.htb www.inlanefreight.htb forum.inlanefreight.htb
Built from HTB Academy's Information Gathering – Web Edition module — labs, lessons and the CWES exam are on HTB Academy.