summaryrefslogtreecommitdiff
path: root/htroot/CrawlStartExpert.html
AgeCommit message (Expand)Author
2026-07-05removed the snapshot feature and all dependencies from non-java codeMichael Peter Christen
2025-08-04more prominent place for collections in crawl start to promote correctMichael Peter Christen
2023-01-16added canonical filterMichael Peter Christen
2023-01-15front-end integration of tag valencyMichael Peter Christen
2022-09-29new link to crawlstart api documentationMichael Peter Christen
2021-08-03fixed doku linkMichael Peter Christen
2020-12-19enabling all crawl profiles in all network modesMichael Peter Christen
2020-01-16enhanced crawl start url check experienceMichael Peter Christen
2019-05-01New optional crawl filter on the URL a doc must match to crawl its linksluccioman
2018-10-25Added new crawler attribute for finer control over Media Type detectionluccioman
2018-10-16Added a crawl start hint message on availability or not of wkhtmltopdfluccioman
2018-07-06Added and updated hint messages about remote crawler statusluccioman
2018-06-19Added a new crawler document filter type using Solr syntaxluccioman
2018-03-23Added a crawl filtering possibility on documents Media Type (MIME)luccioman
2018-03-10added nav filterMichael Peter Christen
2018-02-16Fixed CrawlStartExpert.html HTML validation errorsluccioman
2018-02-16Issue #156 : new option to clean up (or not) search cache on crawl startluccioman
2018-02-10Fixed issue #158 : completed div CSS class ignore in crawlluccioman
2017-12-19Updated links to Java Regular Expressions documentation to version 8luccioman
2017-12-09added a crawl filter based on <div> tag class namesMichael Peter Christen
2017-06-17Limit the number of initially previewed links in crawl start pages.luccioman
2016-11-14Updated Pattern JavaDoc links to current minimum (1.7) JDK version.luccioman
2016-11-12Converted one more set of URLs to pure relative ones.luccioman
2015-05-08added must-not-match filter to snapshot generation.Michael Peter Christen
2015-04-15enhanced timezone managament for indexed data:Michael Peter Christen
2015-02-04remove remote indexing option in crawl start if not in p2p modeMichael Peter Christen
2015-01-30added a html field scraper which reads text from html entities of aMichael Peter Christen
2014-12-09enhanced the snapshot functionality:Michael Peter Christen
2014-12-02get cloned crawl start parameter for snapshotsMichael Peter Christen
2014-12-01YaCy can now create web page snapshots as pdf documents which can laterMichael Peter Christen
2014-08-27added hint to the regular expression testerorbiter
2014-07-18added an option to set 'obey nofollow' for links with rel="nofollow"Michael Peter Christen
2014-06-27fixed external linkMichael Peter Christen
2014-05-17fix: allow enable of CrawlStartExpert.html #filereger
2014-04-30use submitted default userAgent if cloning a crawlMichael Peter Christen
2014-03-31made crawl start pages public since they do not reveal individualorbiter