All articles

Information Gathering — Web Reconnaissance, DNS, Subdomains and Fingerprinting (HTB CWES)

Built from the Information Gathering – Web Edition module of the HTB Academy Web Penetration Tester path (HTB CWES). Revision notes: every tool and command kept, the concepts condensed, a full cheatsheet at the end. Labs and lessons live on HTB Academy.

Reconnaissance is the foundation of a web assessment — map the target before touching it. The goal is to uncover assets (pages, subdomains, IPs, technologies), hidden information (backups, config files, comments), the attack surface, and intelligence (people, emails) you can turn into an entry point.

1. Active vs passive recon

Active Passive
How Directly interact with the target Use only public, third-party sources
Examples Port scan, vuln scan, banner grab, vhost/subdomain brute force, crawling WHOIS, DNS lookups, CT logs, search engines, Wayback, social media, GitHub
Detection Logged, can trip IDS/WAF Very low — looks like normal internet use

Start passive to stay quiet, then go active where you need depth. Always have authorization before active recon.

2. WHOIS

WHOIS is the internet's phonebook: who registered a domain (or owns an IP block / ASN). Install and query:

sudo apt install whois -y
whois inlanefreight.com

A record yields the registrar, registrant/admin/tech contacts, creation/expiry dates, and name servers. For a pentester that's social-engineering fodder (names, emails, phones), infrastructure clues (name servers → hosting), and — via history services like WhoisFreaks — ownership changes over time. Red flags on a suspicious domain: very recent registration, privacy-masked registrant, bulletproof-host name servers.

3. DNS fundamentals

DNS turns names into IPs through a recursive chain of servers.

Your computer / browser needs the IP for www.example.com DNS resolver (ISP) cache miss → starts recursion Root name server points to the .com TLD server TLD name server · .com points to the authoritative server Authoritative name server returns the A record — the IP Your computer connects to the returned IP address
Recursive resolution: resolver → root → TLD → authoritative, which returns the IP.

A zone is a slice of the namespace an authoritative server manages; its zone file holds the records. The hosts file (/etc/hosts on Linux/macOS, C:\Windows\System32\drivers\etc\hosts on Windows) overrides DNS locally — essential for reaching lab vhosts that have no public record:

<IP Address>  <Hostname>  [<Alias> ...]
10.129.42.190  inlanefreight.htb  www.inlanefreight.htb

Record types to know:

Type Meaning
A / AAAA Hostname → IPv4 / IPv6
CNAME Alias → another hostname
MX Mail server(s) for the domain
NS Authoritative name server(s)
TXT Arbitrary text (SPF, verification, _1password=… leaks)
SOA Zone admin info (primary NS, serial, timers)
SRV Host + port for a service
PTR Reverse lookup, IP → hostname

Why it matters: records reveal subdomains, mail servers and infrastructure; a CNAME to a dead host hints at subdomain takeover; a new subdomain appearing over time is a new entry point; TXT records leak which SaaS the org uses.

4. Digging DNS with dig

dig (Domain Information Groper) is the workhorse. Companions: nslookup, host (simpler); dnsenum, fierce, dnsrecon, theHarvester (automated).

dig domain.com            # default A lookup
dig domain.com MX         # mail servers
dig domain.com NS         # name servers
dig domain.com TXT        # TXT records
dig domain.com SOA        # zone admin record
dig @1.1.1.1 domain.com   # query a specific resolver
dig +trace domain.com     # full resolution path (root → authoritative)
dig -x 192.168.1.1        # reverse lookup (PTR)
dig +short domain.com      # just the answer
dig +noall +answer domain.com   # only the answer section

Read the output in four parts: header (status NOERROR, flags), question, answer (name TTL IN A IP), footer (query time, server). ANY is mostly ignored now (RFC 8482). Mind rate limits — excessive queries can be blocked.

5. Subdomain enumeration

Subdomains hide dev/staging boxes, admin portals, legacy apps and leaked data. Two approaches:

Active — zone transfer (AXFR). A misconfigured name server will hand you the entire zone (every subdomain + IP). Rare today, always worth a try:

dig axfr @nsztm1.digi.ninja zonetransfer.me

Active — brute force. Test a wordlist of names against the domain. Tools: dnsenum, fierce, dnsrecon, amass, assetfinder, puredns, ffuf, gobuster.

dnsenum --enum inlanefreight.com \
  -f /usr/share/seclists/Discovery/DNS/subdomains-top1million-20000.txt -r

(-f wordlist, -r recursive.)

Passive — Certificate Transparency logs. Every SSL/TLS cert a CA issues is logged publicly; the SAN field lists subdomains, including dead/old ones. No wordlist, no guessing. Query crt.sh from the CLI:

curl -s "https://crt.sh/?q=facebook.com&output=json" \
  | jq -r '.[] | select(.name_value | contains("dev")) | .name_value' | sort -u

(Censys is the heavier, filterable alternative.)

6. Virtual hosts

One IP can serve many sites; the web server picks which by the HTTP Host header. Non-public vhosts have no DNS record but still respond if you send the right Host — so vhost fuzzing finds them. gobuster (or ffuf/feroxbuster):

gobuster vhost -u http://inlanefreight.htb:81 \
  -w /usr/share/seclists/Discovery/DNS/subdomains-top1million-110000.txt --append-domain

--append-domain is required in current Gobuster; add -t (threads), -k (ignore TLS errors), -o (save output). Note the difference from subdomains: vhost discovery works by varying the Host header against a known IP, not by DNS.

7. Fingerprinting

Identify the stack (server, OS, CMS, framework, WAF) so you can aim exploits precisely.

  • Banner grabbing / headers — the Server and X-Powered-By headers, redirects, and paths like /wp-json/ (WordPress): bash curl -I inlanefreight.com # HTTP headers only curl -I https://www.inlanefreight.com
  • WAF detection — before probing hard, know what's filtering you: bash wafw00f inlanefreight.com # e.g. "behind Wordfence (Defiant) WAF"
  • Nikto — server scanner; fingerprint-only tuning: bash nikto -h inlanefreight.com -Tuning b
  • Others: whatweb, nmap (service/OS, -sV/-O, NSE scripts), and the GUI/online Wappalyzer, BuiltWith, Netcraft.

8. Crawling and hidden paths

Crawling (spidering) follows links from a seed page to map the site and harvest links, comments, metadata and sensitive files (.bak, .old, web.config, settings.php, logs, keys). Value comes from context — a /files/ directory in the links plus a "file server" comment is a lead.

robots.txt (site root) lists paths the owner wants crawlers to skip — which is exactly where the interesting stuff often is:

User-agent: *
Disallow: /admin/
Disallow: /private/
Allow:    /public/
Sitemap:  https://example.com/sitemap.xml

Read Disallow entries as a map of hidden directories (admin panels, backups).

.well-known/ URIs (RFC 8615) are standardized metadata endpoints — an IANA registry lists them. security.txt gives a security contact; openid-configuration dumps a JSON of OAuth/OIDC endpoints (authorization_endpoint, token_endpoint, jwks_uri, scopes) — a ready map of the auth surface:

curl -s https://example.com/.well-known/openid-configuration | jq

Automated crawlers: Burp Suite Spider, OWASP ZAP, Scrapy (Python framework), Apache Nutch. The module ships a custom Scrapy spider, ReconSpider, that outputs results.json (emails, links, external/JS files, images, comments):

pip3 install scrapy
python3 ReconSpider.py http://inlanefreight.com   # → results.json

9. Search engine discovery (Google dorking)

Search operators turn Google into a recon tool (the Google Hacking Database has thousands of ready dorks):

Operator Finds
site: Pages on one domain
inurl: / allinurl: Term(s) in the URL
filetype: / ext: Files of a type (pdf, sql, conf)
intitle: / intext: Term in the title / body
cache: Google's cached copy
"…" · - · OR · * · .. Exact phrase · exclude · either · wildcard · number range
site:example.com inurl:login            # login pages
site:example.com filetype:pdf           # exposed documents
site:example.com inurl:config.php       # config files
site:example.com filetype:sql           # database backups
site:example.com (inurl:admin OR inurl:login)

10. Web archives — the Wayback Machine

web.archive.org stores historical snapshots of sites since 1996. Passive and powerful: recover old pages, directories, files and subdomains that are gone from the live site (but maybe still reachable), diff how the site changed, and gather OSINT — all without touching the target.

11. Automating recon

Frameworks chain these tasks for speed, scale and consistency:

  • FinalRecon — headers, WHOIS, SSL, crawl, DNS, subdomains, directories, Wayback, port scan.
  • Recon-ng — modular framework (DNS, subdomains, ports, crawling, exploits).
  • theHarvester — emails, subdomains, hosts, names, ports from public sources.
  • SpiderFoot — broad OSINT automation across many data sources.
  • OSINT Framework — a directory of sources/tools.
git clone https://github.com/thewhiteh4t/FinalRecon.git && cd FinalRecon
pip3 install -r requirements.txt && chmod +x ./finalrecon.py
./finalrecon.py --headers --whois --url http://inlanefreight.com
./finalrecon.py --full --url http://inlanefreight.com     # everything

12. What to carry into the CWES exam

  • Passive first, then active. WHOIS, CT logs (crt.sh), Google dorks and Wayback cost nothing and tip nobody off; brute force and crawling come after.
  • Enumerate subdomains and vhosts — they're different: subdomains via DNS/CT logs, vhosts via the Host header against the IP. Add every host you find to /etc/hosts.
  • Fingerprint before you attack — curl -I, wafw00f, whatweb/nikto. Knowing it's WordPress behind Wordfence changes everything you do next.
  • Mine robots.txt, .well-known and the Wayback Machine for hidden paths and the auth surface.
  • The skills assessment = whois + robots.txt analysis + subdomain brute force + crawl, with discovered subdomains added to your hosts file. Practice the chain end to end.

Cheatsheet — Information Gathering (Web)

WHOIS

sudo apt install whois -y
whois inlanefreight.com

DNS / dig

dig domain.com                 # A record
dig domain.com MX|NS|TXT|SOA|CNAME|AAAA
dig @1.1.1.1 domain.com        # specific resolver
dig +short domain.com          # answer only
dig +noall +answer domain.com  # answer section only
dig +trace domain.com          # full resolution path
dig -x 1.2.3.4                 # reverse (PTR)
dig axfr @ns.target.tld target.tld   # zone transfer (AXFR)
nslookup domain.com            # host domain.com   (simpler tools)

Subdomains

dnsenum --enum DOMAIN -f /usr/share/seclists/Discovery/DNS/subdomains-top1million-20000.txt -r
# CT logs via crt.sh:
curl -s "https://crt.sh/?q=DOMAIN&output=json" | jq -r '.[].name_value' | sort -u
# others: fierce, dnsrecon, amass, assetfinder, puredns, ffuf, gobuster dns

Virtual hosts

gobuster vhost -u http://TARGET_IP:PORT -w WORDLIST --append-domain [-t 50] [-k] [-o out.txt]

Fingerprinting

curl -I https://TARGET            # headers / Server banner
wafw00f TARGET                    # detect WAF
nikto -h TARGET -Tuning b         # fingerprint scan
whatweb TARGET                    # tech fingerprint
nmap -sV -O TARGET                # service/OS
# GUI/online: Wappalyzer, BuiltWith, Netcraft

Crawling / hidden paths

curl -s https://TARGET/robots.txt
curl -s https://TARGET/.well-known/security.txt
curl -s https://TARGET/.well-known/openid-configuration | jq
pip3 install scrapy && python3 ReconSpider.py http://TARGET   # → results.json
# crawlers: Burp Spider, OWASP ZAP, Scrapy, Apache Nutch

Google dorks

site:TARGET inurl:login
site:TARGET filetype:pdf          (also: filetype:sql | ext:conf | ext:cnf)
site:TARGET inurl:config.php
intitle:"index of"  |  intext:"password"  |  cache:TARGET
# Google Hacking Database: exploit-db.com/google-hacking-database

Archives & automation

# Wayback Machine: https://web.archive.org/
git clone https://github.com/thewhiteh4t/FinalRecon.git && cd FinalRecon
pip3 install -r requirements.txt && chmod +x ./finalrecon.py
./finalrecon.py --headers --whois --url http://TARGET
./finalrecon.py --full --url http://TARGET
# frameworks: Recon-ng, theHarvester, SpiderFoot, OSINT Framework

/etc/hosts (add discovered hosts)

<TARGET_IP>  inlanefreight.htb  www.inlanefreight.htb  forum.inlanefreight.htb

Built from HTB Academy's Information Gathering – Web Edition module — labs, lessons and the CWES exam are on HTB Academy.