I just noticed that Googlebot has begun crawling the anonymous FTP here. :-D
To my surprise, they seem to try to enter a /robots.txt subdirectory before
trying to download /robots.txt:
To my surprise, they seem to try to enter a /robots.txt subdirectory
that's how they decide if it is a file or a directory... if it is a
file, then they download it...
Yes, I suspected that might be the case... I wonder, though, if they
do anything special if the cd succeeds
or if they just conclude "not a file" and skip it. (According to the
spec, they shouldn't look for anything inside that directory, if it exists.)
http://example.com/folder/robots.txt
"Not a valid robots.txt file. Crawlers don't check for robots.txt files in subdirectories. "
| Sysop: | Winzlo |
|---|---|
| Location: | Minnesota, USA |
| Users: | 11 |
| Nodes: | 16 (0 / 16) |
| Uptime: | 495940:14:26 |
| Calls: | 82 |
| Files: | 1,070 |
| D/L today: |
27 files (11,920K bytes) |
| Messages: | 286,975 |