User-agent: * Allow: / Disallow: /shop/admin/ Disallow: /shop/admin Disallow: /shop/admin/goods/origin_check_frontpage.php Disallow: /shop/goods/indb_err_img.php Disallow: /m2/myp/ Disallow: /m2/myp/review_register.php Disallow: /m2/myp/goods_qna_register.php Disallow: /m2/goods/review_register.php Disallow: /m2/goods/goods_qna_register.php Disallow: /shop/goods/proxy_changer2.php Disallow: /shop/goods/goods_qna_list.php Disallow: /shop/goods/goods_review_viewimg.php Disallow: /shop/goods/popup_request_stocked_noti.php Disallow: /shop/goods/goods_qna_register.php Disallow: /m2/goods/goods_qna_list.php Disallow: /shop/goods/goods_qna_register.php Disallow: /shop/goods/goods_review_list.php # --- 2026-08 SEO / REVERTED 2026-08-28 --- # These two handlers were disallowed because they ate ~60% of Naver Yeti's # crawl budget (12,371 hits, only 1,693 on product pages). # # Blocking them backfired. Naver's site diagnosis type "access-restricted # resource" went 0 -> 381 on 08-27~28. Its renderer executes the product # page's XHR, hits the block, and downgrades the document to "cannot be # recognised as a normal document". Every other diagnosis type was falling # in that window; only this one spiked, and the URLs it named are exactly # these two. Losing the indexing verdict costs more than the crawl budget # saves, so they are open again. # # Instead both handlers now send "X-Robots-Tag: noindex" in their own # response, so they stay crawlable (renderable) but are never indexed: # shop/proc/indb.cart.tab.php # m2/proc/mAjaxAction.php # # Do NOT re-add these two Disallow lines. If crawl budget becomes a problem # again, rate-limit or cache the handlers instead of blocking them. # --- 2026-08-26 : search result pages --- # Bingbot crawled 6,396 search URLs in 2.75 days. That was 100% of all # search crawling here, and every one was served a 302. # The queries are spam injected from outside (Chinese spam text, # +51cg365.com, script fragments): someone links to our search so that # their text renders on our domain. Search results have no index value. # Category listings use "category=" and are NOT matched by these rules. Disallow: /shop/goods/goods_search.php Disallow: /*sword= Disallow: /*searched= # --- 2026-08-28 : mobile search --- # The rules above only catch the PC search ("sword="). Mobile search uses # "kw=" (/m2/goods/list.php?kw=..) and was still crawled: 3,299 bot requests # in the 49 days to 2026-08-27, 2,808 of them search result lists. # Search results also link products as /m2/goods/view.php?kw=..&goodsno=.. # Those are duplicates - every product page has rel=canonical to the PC URL, # and the sitemaps contain no kw= URL at all. # # The Allow line matters. Of 580 distinct kw= URLs bots actually crawled, # 417 carry an EMPTY kw= tacked onto a category listing, a product page or # an old URL that 301s to the canonical (deliberate link recovery, 2026-08-26). # All 417 end at "kw=", so "$" hands them back. Blocking those would have # shut off category listings and killed the 301 recovery. # # Do NOT write this as "Disallow: /*&kw=" - verified with a Google-spec parser # (protego), that form matches nothing at all. "/*kw=" is what works, and no # other parameter on this site ends in "kw". Disallow: /*kw= Allow: /*kw=$ # --- 2026-08-28 : third-party site audit crawler --- # AhrefsSiteAudit made 93,203 requests from 1,229 IPs in a single hour # (2026-08-28 00:00 KST). The shop owner did not run it, so this is an # outside party auditing our site. It brings no visitors, and a burst that # size is heavy for the server. Ahrefs honours robots.txt for this agent. User-agent: AhrefsSiteAudit Disallow: / # --- 2026-08-28 : SEO tools and scrapers that send no visitors --- # Measured over the 49 days to 2026-08-28: # SemrushBot 55,254 / Bytespider 48,031 / DataForSeoBot 32,683 / # SemrushBot-SA 512 / MJ12bot 432 / PetalBot 371 = about 137,000 requests. # None of them referred a single visitor in that window. They only eat crawl # budget and server capacity. AhrefsBot is added for the same reason as # AhrefsSiteAudit above. # # AI crawlers are deliberately NOT blocked. GPTBot / OAI-SearchBot / # ClaudeBot / PerplexityBot do send real people: ChatGPT alone sent 511 # visits in 49 days and doubled month over month (425/day in July -> # 868/day in August). Search engines (Googlebot, Yeti, bingbot, Daum, # Applebot, Yandex) stay allowed too. # # Note: Bytespider has a reputation for ignoring robots.txt. If its request # count does not drop within a few days, block it by User-Agent in .htaccess # instead - but check the log first, that rule is riskier. User-agent: AhrefsBot User-agent: SemrushBot User-agent: SemrushBot-SA User-agent: SiteAuditBot User-agent: Bytespider User-agent: DataForSeoBot User-agent: MJ12bot User-agent: PetalBot Disallow: / Sitemap: https://www.rcbank.co.kr/sitemap.xml