"Inference engine setup" "Model assignment preview" "Index creation" "RAG configuration" "Tools configuration" "Log report monitor" "Shield definition" AI Lab Build System Craft your AI toolkit Complete the quests below to unlock YaCy's AI sidekick: bind an inference engine, load production models, feed it with your index, then wire RAG and shields. 0 / 6 unlocked Mandatory Needs setup Bind an inference engine Pick your host (Ollama, LM Studio, OpenAI-compatible) and give YaCy a place to send prompts. Open engine setup Set hoststub, API keys, and defaults to unlock downloads. Populate the Production Models Matrix Assign models for chat, search, translation, and more. This is your loadout bench. Go to Production Models Matrix Deploy at least one model, then assign capabilities (chat, search-query, tooling, vision). Optional Grow a search index Create a local index for grounding: crawl a site or import a pack to give your AI facts to cite. Start a crawl Import an index pack Indexed documents: required to unlock (need at least 1000 documents). Wire RAG retrieval Map which production models answer search-query and Q/A pairs so the RAG proxy can mix search with chat. Wire RAG prompts Test in Chat Set the search-query and qapairs columns to connect retrieval to your chat flow. Enable/Disable Tools Superpowers for the YaCy Chat Open tools configuration Tune descriptions and set maxCallsPerTurn per tool (0 disables a tool). Monitor log reports Assign a log-report model, then review generated hourly and daily self-enhancement reports. Open log reports Assign log-report model Report generation stays inactive until a production model is assigned to the log-report role. Define a shield Add guardrails: access rates, grant or deny non-localhost access. Activate the front page link for chat to complete this quest. Open shield settings Store your shield directives (system prompts, stop words) as properties, then exercise them in chat. Wire RAG Retrieval Shield Control who can access the chat interface and rate-limit non-localhost clients to protect your peer and LLM backends from overload. Overall Load Protection Recent access volume across all clients (localhost included). You can enforce global limits here to protect the host. Requests / minute Requests / hour Requests / day Limit for all requests, including localhost Per minute: Per hour: Per day: Guest Access Control & Rate Limits By default only localhost may reach the chat UI. Enable non-localhost access and throttle requests to reduce abuse. Allow non-localhost clients to access the chat interface Requests from non-localhost will be throttled using these caps: Front Page Link Expose a shortcut to the chat UI on the search front page if you want users to discover it. Show a link to yacychat.html on the search front page Save Shield Settings "YaCy Access Grid" Server Access Grid This images shows incoming connections to your YaCy peer and outgoing connections from your peer to other peers and web servers Server Access Overview Host Access Count During last Second last Minute last 10 Minutes last Hour The following hosts are registered as source for brute-force requests to protected pages Access Times Server Access Details This is a list of requests (max. 1000) to the local http server within the last hour. Date Path Local Search Log This is a list of searches that had been requested from this' peer search interface Requesting Host Offset Expected Results Returned Results Known Results Used Time (ms) URL fetch (ms) Snippet comp (ms) Query User Agent Top Search Words (last 7 Days) Local Search Host Tracker Count Queries Per Last Hour Access Dates Remote Search Log This is a list of searches that had been requested from remote peer search interface Peer Name Search Word Hashes Remote Search Host Tracker "Save" Autocrawler Autocrawler automatically selects and adds tasks to the local crawl queue. This will work best when there are already quite a few domains in the index. Autocralwer Configuration You need to restart for some settings to be applied Enable Autocrawler: Deep crawl every Nth document: Warning: if this is bigger than "Rows to fetch" only shallow crawls will run. Rows to fetch at once: Recrawl only older than # days: Get hosts by query: Can be any valid Solr query. Shallow crawl depth (0 to 2): Deep crawl depth (1 to 5): Index text: Index media: "API" "no previous page" "previous page" "no next page" "next page" "Apply edited next execution dates" "clone" "yyyy/MM/dd HH:mm:ss" "Execute Selected Actions" "Delete Selected Actions" "Delete all Actions which had been created before " Process Automation This table shows actions that had been issued on the YaCy interface. These recorded actions can be used to repeat specific actions and to send them to a scheduler for a periodic execution. The information that is presented on this page can also be retrieved as XML. Click the API icon to see the XML. Recorded Actions Type Comment Call Count Recording&nbsp;Date Last&nbsp;Exec&nbsp;Date Next&nbsp;Exec&nbsp;Date Apply Event Trigger Scheduler URL no event activate event off run once run regular after start-up at 00:00h at 01:00h at 02:00h at 03:00h at 04:00h at 05:00h at 06:00h at 07:00h at 08:00h at 09:00h at 10:00h at 11:00h at 12:00h at 13:00h at 14:00h at 15:00h at 16:00h at 17:00h at 18:00h at 19:00h at 20:00h at 21:00h at 22:00h at 23:00h no repetition activate scheduler minutes hours days 1 day 2 days 3 days 4 days 5 days 6 days 1 week 2 weeks 3 weeks 1 month 2 months 3 months 6 months 9 months 1 year 2 years Result of API execution Status "Check" "Change Selected" "Delete Selected" Blacklist Cleaner Here you can remove or edit illegal or double blacklist-entries. Check list Allow regular expressions in host part of blacklist entries. The blacklist-cleaner only works for the following blacklist-engines up to now: Two wildcards in host-part Either subdomain or wildcard Path is invalid Regex Wildcard not on begin or end Host contains illegal chars Double Host is invalid Regex No Blacklist selected "Load new blacklist items" "Export list as XML" "Export list as text" Blacklist Import Used Blacklist engine: Import blacklist items from... other YaCy peers: URL: plain text file: Upload a regular text file which contains one blacklist entry per line. XML file: Upload an XML file which contains one or more blacklists. Export blacklist items to... Here you can export a blacklist as an XML file. This file will contain additional information about which cases a blacklist is activated for. all Here you can export a blacklist as a regular text file with one blacklist entry per line. This file will not contain any additional information. "Test" Blacklist Test Used Blacklist engine: Test list: It is blocked for the following cases: is not blocked Crawling DHT News Proxy Search Surftips The tested URL was not valid. "create" "Add URL pattern" "set" "Save URL pattern(s)" "Share/don't share this list" "Delete this list" "Save" Blacklist Administration This function provides an URL filter to the proxy; any blacklisted URL is blocked from being loaded. You can define several blacklists and activate them separately. You may also provide your blacklist to other peers by sharing them; in return you may collect blacklist entries from other peers. Active list: No blacklist selected Select list to edit: not shared shared Create new list: A legal name is made up from a letter, digit, minus, plus or underscore as the first character followed by letters, digits, minus, plus, underscores or dots. An error occurred while moving entries to the target list. Add new pattern: domain.net/fullpath domain.net/* sub.domain.*/* domain.*/* Blacklist Pattern Edit selected pattern(s) Delete selected pattern(s) Move selected pattern(s) to Show entries: Entries per page: Edit existing pattern(s): An error occurred while editing the following entries. Please check syntax. Activate this list for ... "RSS" "Submit" "Preview" "Discard" "Yes, delete it." "No, leave it." "Import" &lt;&lt; previous entries next entries &gt;&gt; Blog-Home Edit Author: Subject: Text: Comments: deactivated activated moderated Preview No changes have been submitted so far! Access denied To edit or create blog-entries you need to be logged in as Admin or User who has Blog rights. Are you sure... Confirm deletion XML-Import Import was successful! Import failed, maybe the supplied file was no valid blog-backup? Please select the XML-file you want to import: "Submit" "Preview" "Discard" Blog-Home Comments: &lt;&lt; previous entries next entries &gt;&gt; Comments are not allowed for this posting! Comment on this Blog Author: Subject: Text: "RSS" "create" "Save" "import" "API" "start it" "stop it" "private bookmark" "public bookmark" Bookmarks Login List Bookmarks Add Bookmark Import Bookmarks Bookmarks (XBEL) Bookmarks (XML) Bookmarks (RSS) Edit Bookmark URL: Title: Description: Query: Folder (/folder/subfolder): Tags (comma separated): Public: yes no Bookmark is a newsfeed Import XML Bookmarks File: import as Public: Import HTML Bookmarks Default Tags: The bookmarks list can also be retrieved as RSS feed. This can also be done when you select a specific tag. Click the API icon to load the RSS from the current selection. Folders Bookmark Folder Tags Auto Search start autosearch of new bookmarks autosearch queue: received results: current query: This starts a search of new or modified bookmarks since startup in folder "search" with "query=&lt;original_search_term&gt;" Every peer online will be ask for results. Bookmark List Tagged with | Edit Delete Info search previous page next page Show Bookmarks per page. Image Collage Private Queue Public Queue User List User Accounts User First name Last name Address Last Access Rights Time Traffic "Define Administrator" "Set Access Rules" "Edit User" "Delete User" "Save User" User Administration Generic error. Passwords do not match. Username too short. Username must be &gt;= 4 Characters. Username already used (not allowed). <b>WARNING</b> This YaCy instance can be administered with the account "admin" and the default password "yacy". Change the password as soon as possible! Admin Account Access from localhost without account Access to your peer from your own computer (localhost access) is granted with administrator rights. No need to configure an administration account. This setting is convenient but less secure than using a qualified admin account. Please use with care, notably when you browse untrusted and potentially malicious websites while running your YaCy peer on the same computer. Access only with qualified account This is required if you want a remote access to your peer, but it also hardens access controls on administration operations of your peer. Peer User: New Peer Password: Repeat Peer Password: Access Rules Protection of all pages: if set to on, access to all pages need authorization; if off, only pages with "_p" extension are protected. User Accounts Select user New user Username Password Repeat password First name Last name Address Rights: Timelimit Time used "Use" "Delete" "Set Colors" "Install" Appearance and Integration You can change the appearance of the YaCy interface with skins. The selected skin and language also affects the appearance of the search page. change the appearance of the search page here. Skin Selection Select one of the default skins. <b>After selection it might be required to reload the web page while holding the shift key to refresh cached style files.</b> Current skin Available Skins Skin Color Definition The generic skin 'generic_pd' can be configured here with custom colors: Background Text Legend Table&nbsp;Header Table&nbsp;Item Table&nbsp;Item&nbsp;2 Table&nbsp;Bottom Border&nbsp;Line Sign&nbsp;'bad' Sign&nbsp;'good' Sign&nbsp;'other' Search&nbsp;Headline Search&nbsp;URL Search&nbsp;URL&nbsp;+&nbsp;hover Skin Download Skins can be installed from download locations: Install new skin from URL Use this skin Make sure that you only download data from trustworthy sources. The new Skin file might overwrite existing data if a file of the same name exists already. Error saving the skin. "ok" "Use the browser preferred language if available" "Click to generate translated pages" "Active : translated pages are available" "Usecase Freeworld" "Usecase Portal" "Usecase Intranet" "warning" "Set Configuration" Basic Configuration Your port has changed. Please wait 10 seconds. <b>WARNING</b> This YaCy instance can be administered with the account "admin" and the default password "yacy". Your YaCy Peer needs some basic information to operate properly Select a language for the interface: Browser English Deutsch Fran&ccedil;ais Greek Italiano Español Use Case: what do you want to do with YaCy: Can not leave from Intranet Indexing : one or more remote Solr instances are attached and may contain private documents indexed. One or more remote Solr instances are attached and may contain indexed public documents irrelevant to your local domain. One or more remote Solr instances are attached. Community-based web search Search portal for your own web pages Intranet Indexing Join and support the global network 'freeworld', search the web with an uncensored user-owned search network Your YaCy installation behaves independently from other peers and you define your own web index by starting your own web crawl. This can be used to search your own web pages or to define a topic-oriented search portal. Create a search portal for your intranet or web pages or your (shared) file system. URLs may be used with http/https/ftp and a local domain name or IP, or with an URL of the form file:///&lt;path&gt; or smb://&lt;server&gt;/&lt;path&gt; Your peer name has not been customized; please set your own peer name You may change your peer name Peer Name: Your peer can be reached by other peers Peer Port: with SSL (https enabled Configure your router for YaCy using UPnP: Configuration was not successful. This may take a moment. Your Browser will reload the YaCy UI with the new port in 5 seconds... What you should do next: Your basic configuration is complete! You can now (for example): Your Peer name is a default name; please set an individual peer name. You did not open a port in your firewall or your router does not forward the server port to your peer. This is needed if you want to fully participate in the YaCy network. You can also use your peer without opening it, but this is not recommended. "A cache hit occurs when the requested data can be found in a cache." "Concurrent access timeout info" "Set" "Delete" Hypertext Cache Configuration The HTCache stores content retrieved by the HTTP and FTP protocol. Documents from smb:// and file:// locations are not cached. The cache is a rotating cache: if it is full, then the oldest entries are deleted and new one can fill the space. HTCache Configuration Cache hits The path where the cache is stored The current size of the cache The maximum size of the cache MB Compression level Concurrent access timeout The maximum time to wait for acquiring a synchronization lock on concurrent get/store cache operations. Beyond this limit, the crawler or proxy falls back to regular remote resource loading. milliseconds Cleanup Cache Deletion Delete HTTP &amp; FTP Cache Delete robots.txt Cache "heuristic:&lt;name&gt; (redundant)" "heuristic:&lt;name&gt; (new link)" "add" "Save" "reset to default list" "discover from index" "switch Solr fields on" Heuristics Configuration When a search heuristic is used, the resulting links are not used directly as search result but the loaded pages are indexed and stored like other content. This ensures that blacklists can be used and that the searched word actually appears on the page that was discovered by the heuristic. The success of heuristics are marked with an image ( ) below the favicon left from the search result entry: The search result was discovered by a heuristic, but the link was already known by YaCy The search result was discovered by a heuristic, not previously known by YaCy 'site'-operator: instant shallow crawl When a search is made using a 'site'-operator (like: 'download site:yacy.net') then the host of the site-operator is instantly crawled with a host-restricted depth-1 crawl. That means: right after the search request the portal page of the host is loaded and every page that is linked on this page that points to a page on the same host. Because this 'instant crawl' must obey the robots.txt and a minimum access time for two consecutive pages, this heuristic is rather slow, but may discover all wanted search results using a second search (after a small pause of some seconds). search-result: shallow crawl on all displayed search results add as global crawl job When a search is made then all displayed result links are crawled with a depth-1 crawl. This means: right after the search request every page is loaded and every page that is linked on this page. If you check 'add as global crawl job' the pages to be crawled are added to the global crawl queue (remote peers can pickup pages to be crawled). Default is to add the links to the local crawl queue (your peer crawls the linked pages). opensearch load external search result list from active systems below When using this heuristic, then every new search request line is used for a call to listed opensearch systems. 20 results are taken from remote system and loaded simultaneously, parsed and indexed immediately. Available/Active Opensearch System Active Title Comment Url delete new With the button "discover from index" you can search within the metadata of your local index (Web Structure Index) to find systems which support the Opensearch specification. The task is started in the background. It may take some minutes before new entries appear (after refreshing the page). "Use" "Delete" "Install" Language selection You can change the language of the YaCy-webinterface with translation files. Current language default(english) Author(s) (chronological) Send additions to maintainer Available Languages Download Language File Supported formats are the internal language file (extension .lng) or XLIFF (extension .xlf) format. Install new language from URL Use this language Make sure that you only download data from trustworthy sources. The new language file might overwrite existing data if a file of the same name exists already. Error saving the language file. "Change Network" "Save" "Transport Layer Security" "Secure Sockets Layer" Network Configuration Accepted Changes. Inapplicable Setting Combination: No changes were made! For P2P operation, at least DHT distribution or DHT receive (or both) must be set. You have thus defined a Robinson configuration. Global Search in P2P configuration is only allowed, if index receive is switched on. You have a P2P configuration, but are not allowed to search other peers. For Robinson Mode, index distribution and receive is switched off. Network and Domain Specification YaCy can operate a computing grid of YaCy peers or as a stand-alone node. To control that all participants within a web indexing domain have access to the same domain, this network definition must be equal to all members of the same YaCy network. Network Definition Enter custom URL... Remote Network Definition URL Network Nick Long Description Indexing Domain DHT Distributed Computing Network for Domain Enable Peer-to-Peer Mode to participate in the global YaCy network, or if you want your own separate search cluster with or without connection to the global network. Enable 'Robinson Mode' for a completely independent search engine instance, without any data exchange between your peer and other peers. Peer-to-Peer Mode Index Distribution This enables automated, DHT-ruled Index Transmission to other peers. enabled disabled during crawling disabled during indexing Index Receive Accept remote Index Transmissions. This works only if you have a senior peer. The DHT-rules do not work without this function. reject accept transmitted URLs that match your blacklist allow deny remote search Robinson Mode If your peer runs in 'Robinson Mode' you run YaCy as a search engine for your own search portal without data exchange to other peers. There is no index receive and no index distribution between your peer and any other peer. In case of Robinson-clustering there can be acceptance of remote crawl requests from peers of that cluster. Private Peer Your search engine will not contact any other peer, and will reject every request. Public Peer You are visible to other peers and contact them to distribute your presence. Your peer does not accept any outside index data, but responds on all remote search requests. Public Cluster Your peer is part of a public cluster within the YaCy network. Index data is not distributed, but remote crawl requests are distributed and accepted Search requests are spread over all peers of the cluster, and answered from all peers of the cluster. List of .yacy or .yacyh - domains of the cluster: (comma-separated) Peer Tags When you allow access from the YaCy network, your data is recognized using keywords. Please describe your search portal with some keywords (comma-separated). If you leave the field empty, no peer asks your peer. If you fill in a '*', your peer is always asked. Outgoing communications encryption Protocol operations encryption Prefer HTTPS for outgoing connexions to remote peers. When <abbr title="Transport Layer Security">TLS</abbr>/<abbr title="Secure Sockets Layer">SSL</abbr> is enabled on remote peers, it should be used to encrypt outgoing communications with them (for operations such as network presence, index transfer, remote crawl...). Please note that contrary to strict TLS, certificates are not validated against trusted certificate authorities (CA), thus allowing YaCy peers to use self-signed certificates. "Submit" Parser Configuration Content Parser Settings With this settings you can activate or deactivate parsing of additional content-types based on their MIME-types. For a detailed description of the various MIME-types take a look at Extension Mime-Type "Remote results resorting can be triggered once the 'Refresh sorting' button (near the 'Search' button) becomes available." "This usually improves ranking accuracy, but doesn't work well for users who have Javascript disabled, are using screen readers, or are on slow computers." "idea" "Detailed statistics" "Change Search Page" "Set to Default Values" Integration of a Search Portal If you like to integrate YaCy as portal for your web pages, you may want to change icons and messages on the search page. The search page may be customized. You can change the 'corporate identity'-images, the greeting line and a link to a home page that is reached when the 'corporate identity'-images are clicked. Greeting Line URL of Home Page URL of a Small Corporate Image URL of a Large Corporate Image Alternative text for Corporate Images Enable Search for Everyone? Search is available for everyone Only the administrator is allowed to search Show Navigation Bar on Search Page? Show Navigation Top-Menu no link to YaCy Menu (admin must navigate to /Status.html manually) Show Advanced Search Options on Search Page? Show Advanced Search Options on index.html do not show Advanced Search Media Search Extended Strict Control whether media search results are as default strictly limited to indexed documents matching exactly the desired content domain (images, videos or applications specific), or extended to pages including such medias (provide generally more results, but eventually less relevant). Remote results resorting On demand, server-side Automated, with JavaScript in the browser. Automated results resorting with JavaScript makes the browser load the full result set of each search request. This may lead to high system loads on the server. Remote search encryption Prefer https for search queries on remote peers. When SSL/TLS is enabled on remote peers, https should be used to encrypt data exchanged with them when performing peer-to-peer searches. Please note that contrary to strict TLS, certificates are not validated against trusted certificate authorities (CA), thus allowing YaCy peers to use self-signed certificates. Snippet Fetch Strategy &amp; Link Verification Speed up search results with this option! (use CACHEONLY or FALSE to switch off verification) Counts by origin : NOCACHE: no use of web cache, load all snippets online IFFRESH: use the cache if the cache exists and is fresh otherwise load online IFEXIST: use the cache if the cache exist or load online If verification fails, delete index reference CACHEONLY: never go online, use all content from cache. If no cache entry exist, consider content nevertheless as available and show result without snippet FALSE: no link verification and not snippet generation: all search results are valid without verification Greedy Learning Mode Index remote results add remote search results to the local index <b>( default=on, it is recommended to enable this option ! )</b> Limit size of indexed remote results maximum allowed size in kbytes for each remote search result to be added to the local index (for example, a 1000kbytes limit might be useful if you are running YaCy with a low memory setup) Default Pop-Up Page Status Page Search Front Page Search Page (small header) Interactive Search Page Default maximum number of results per page Default index.html Page (by forwarder) Target for Click on Search Results "_blank" (new window) "_self" (same window) "_parent" (the parent frame of a frameset) "_top" (top of all frames) "searchresult" (a default custom page name for search results) Special Target as Exception for an URL-Pattern Pattern: Exclude Hosts List of hosts that shall be excluded from search results by default but can be included using the site:&lt;host&gt; operator: 'About' Column<br/>(shown in a column alongside<br/>with the search result page) (Headline)</br> (Content) The search page can be integrated in your own web pages with an iframe. Simply use the following code: This would look like: For a search page with a small header, use this code: A third option is the interactive search. Use this code: "Save" Your Personal Profile You can create a personal profile here, which can be seen by other YaCy-members Name Nick Name eMail ICQ Jabber Yahoo! MSN Skype Comment "Save" "Clear" Advanced Config Here are all configuration options from YaCy. You can change anything, but some options need a restart, and some options can crash YaCy, if wrong values are used. For explanation please look into defaults/yacy.init "Save restrictions" Exclude Web-Spiders Here you can set up a robots.txt for all webcrawlers that try to access the webinterface of your peer. robots.txt is a voluntary agreement most search-engines (including YaCy) follow. It disallows crawlers to access webpages or even entire domains. Unable to access the local file: Deletion of htroot/robots.txt failed Deny access to Entire Peer Status page Network pages Surftips News pages Blog Wiki Public bookmarks Home Page File Share Impressum "Search" Integration of a Search Box We give information how to integrate a search box on any web page that calls the normal YaCy search window. Simply use the following code: This would look like: MySearch This does not use a style sheet file to make the integration into another web page with a different style sheet easier. You would need to change the following items: Replace the given colors #eeeeee (box background) and #cccccc (box border) Replace the word "MySearch" with your own message "Top navigation bar" "Enable login link/status" "Log in to use extended search features" "You are authenticated as userName" "Help" "Protocols" "Tag cloud" "earthsearchlogo" "Delete navigator" "Sorted by descending counts" "Sorted by ascending counts" "Sorted by descending labels" "Sorted by ascending labels" "search..." "Maximum days number in the histogram. Beware that a large value may trigger high CPU loads both on the server and on the browser with large result sets." "info" "Website favicon" "Last known modification date" "Browse index" "Raw ranking score value" "Date" "Size" "Add navigator" "Save Settings" "Set Default Values" Search Result Page Layout Configuration Below is a generic template of the search result page. Mark the check boxes for features you would like to be displayed. Page Template Toggle navigation Log in userName Search Interfaces<b class="caret"></b> Administration &raquo; http https ftp smb file Tag Topics Cloud Location show search results on map Sort by Descending counts Ascending counts Descending labels Ascending labels Vocabulary search Text Images Audio Video Applications more options Date Navigation Maximum range (in days) Show websites favicon Not showing websites favicon can help you save some CPU time and network bandwidth. Title of Result Description and text snippet of the search result http://url-of-the-search-result.net Tags keyword subject keyword2 keyword3 Max. tags initially displayed (remaining can then be expanded) 42 kbyte Metadata Parser Citation Pictures Cache View via Proxy Ranking: 1.12195955E9 For this option URL proxy must be enabled. menu: System Administration > Advanced Settings Menu: System Administration > Advanced Settings > Debug/Analysis Settings Add Navigators append max. items "Download Release" "Check for new Release" "Install Release" "Delete Release" "Check + Download + Install Release Now" "Submit" System Update Release will be installed. Please wait. This servlet can only be used on operating systems that are currently supported for deploy functions. If you see this message this means that your operation system is not supported. Manual System Update Current installed Release (unsigned) (signed) Downloaded Releases No downloaded releases available for deployment. (no signature) no&nbsp;automated installation on development environments Automatic Update check for new releases, download if available and restart with downloaded release No more recent release found. Omitting update because this is a development environment. Omitting update because an error occurred while trying to deploy the release. Automated System Update manual update no automatic look-up, updates can be made manually using this interface (see options above) automatic update updates are made within fixed cycles: Time between lookup hours Release blacklist (regex on release number strings) Release type only main releases any release including developer releases Signed autoupdate: only accept signed files Accepted Changes. System Update Statistics Last System Lookup never Last Release Download Last Deploy You installed YaCy with a package manager. To update YaCy, use the package manager: manual update:<br/>apt-get update &amp;&amp; apt-get install yacy automatic update: add the following line to /etc/crontab<br/>0 6 * * * root apt-get update &amp;&amp; apt-get -y --force-yes install yacy "Save User" "Delete User" "ConfigAccountList_p.html" User Account Editor Generic error. Passwords do not match. Username too short. Username must be &gt;= 4 Characters. Username already used (not allowed). Username Password Repeat password First name Last name Address Rights: Timelimit Time used back to user list Server Connection Tracking Incoming Connections Protocol Duration Source IP[:Port] Command ID Outgoing Connections Up-Bytes Dest. IP[:Port] "Set" "Re-Set to default" Content Analysis These are document analysis attributes. Double Content Detection Double-Content detection is done using a ranking on a 'unique'-Field, named 'fuzzy_signature_unique_b'. minTokenLen This is the minimum length of a word which shall be considered as element of the signature. Should be either 2 or 3. quantRate The quantRate is a measurement for the number of words that take part in a signature computation. The higher the number, the less words are used for the signature. For minTokenLen = 2 the quantRate value should not be below 0.24; for minTokenLen = 3 the quantRate value must be not below 0.5. "Check database connection" "Export Content to Packs" "Import Dump" Content Integration: Retrieval from phpBB3 Databases It is possible to extract texts directly from mySQL and postgreSQL databases. Each extraction is specific to the data that is hosted in the database. This interface gives you access to the phpBB3 forums software content. If you read from an imported database, here are some hints to get around problems when importing dumps in phpMyAdmin: before importing large database dumps, set the following Line in phpmyadmin/config.inc.php and place your dump file in /tmp (Otherwise it is not possible to upload files larger than 2MB): deselect the partial import flag When an export is started, pack files are generated into DATA/PACKS/load which are automatically fetched by an indexer thread. All indexed pack files are then moved to DATA/PACKS/loaded and can be re-cycled when an index is deleted. <b>The URL stub</b>,<br />like http://forum.yacy-websuche.de<br />this must be the path right in front of '/viewtopic.php?' <b>Type</b> of database<br />(use either 'mysql' or 'pgsql') <b>Host</b> of the database <b>Port</b> of database service<br />(usually 3306 for mySQL) <b>Name of the database</b> on the host <b>Table prefix string</b> for table names <b>User</b> that can access the database <b>Password</b> for the account of that user given above <b>Posts per file</b><br />in exported packs <b>Import a database dump</b>, Posts in database first entry last entry Import successful! "Enable Cookie Monitoring" "Disable Cookie Monitoring" Cookie Monitor: Incoming Cookies This is a list of Cookies that a web server has sent to clients of the YaCy Proxy: Sending Host Date Receiving Client Cookie "Enable Cookie Monitoring" "Disable Cookie Monitoring" Cookie Monitor: Outgoing Cookies This is a list of cookies that browsers using the YaCy proxy sent to webservers: Receiving Host Date Sending Client Cookie "Check given urls" Crawl Check This pages gives you an analysis about the possible success for a web crawl on given addresses. List of possible crawl start URLs Analysis URL Access Robots Crawl-Delay Sitemap Recently started remote crawls in progress Remote crawl start points, crawl is ongoing Start Time Peer Name Start URL Intention/Description Depth Accept '?' URLs no yes Remote crawl start points, finished: "Terminate" "Delete" "Delete finished crawls" "Edit profile" "Submit changes" Crawler Steering Crawl Scheduler Scheduled Crawls can be modified in this table Crawl Profile Editor Crawl profiles hold information about a crawl process that is currently ongoing. Crawl Profile List Crawl Thread Collections Status Depth Must Match Must Not Match Recrawl if older than Domain Counter Content Max Page Per Domain Accept '?' URLs Fill Proxy Cache Local Text Indexing Local Media Indexing Remote Indexing Running Finished no yes Select the profile to edit false true "An illustration how yacy works" "delete all" "del & blacklist" "clear list" "delete" Crawl Results Overview These are monitoring pages for the different indexing queues. YaCy knows 5 different ways to acquire web indexes. The details of these processes (1-5) are described within the submenu's listed above which also will show you a table with indexing results so far. The information in these tables is considered as private, so you need to log-in with your administration password. Case (6) is a monitor of the local receipt-generator, the opposed case of (1). It contains also an indexing result monitor but is not considered private since it shows crawl requests from other peers. Case (7) occurs if pack files are imported The image above illustrates the data flow initiated by web index acquisition. Some processes occur double to document the complex index migration structure. (1) Results of Remote Crawl Receipts This is the list of web pages that this peer initiated to crawl, but had been crawled by <em>other</em> peers. This is the 'mirror'-case of process (6). Every page that a remote peer indexes upon this peer's request is reported back and can be monitored here. No remote crawl results can currently been added to the local index as the remote crawler is disabled on this peer. (2) Results for Result of Search Queries This index transfer was initiated by your peer by doing a search query. The index was crawled and contributed by other peers. <em>Use Case:</em> This list fills up if you do a search query on the 'Search Page' (3) Results for Index Transfer The url fetch was initiated and executed by other peers. These links here have been transmitted to you because your peer is the most appropriate for storage according to the logic of the Global Distributed Hash Table. <em>Use Case:</em> This list may fill if you check the 'Index Receive'-flag on the 'Index Control' page (4) Results for Proxy Indexing These web pages had been indexed as result of your proxy usage. <strong>No personal or protected page is indexed</strong>; such pages are detected by Cookie-Use or POST-Parameters (either in URL or as HTTP protocol) and automatically excluded from indexing. <em>Use Case:</em> You must use YaCy as proxy to fill up this table. Set the proxy settings of your browser to the same port as given on the 'Settings'-page in the 'Proxy and Administration Port' field. (5) Results for Local Crawling These web pages had been crawled by your own crawl task. <em>Use Case:</em> start a crawl by setting a crawl start point on the 'Index Create' page. (6) Results for Global Crawling These pages had been indexed by your peer, but the crawl was initiated by a remote peer. This is the 'mirror'-case of process (1). The remote crawler is currently disabled (7) Results from pack import These records had been imported from pack files in DATA/PACKS/load The stack is empty. Domain URLs Blacklist to use Collection Initiator Executor Modified Words Title Country IP of Host URL no title "API" "info" "empty" "Show all links" "Media Type checking info" "Media Type filter info" "Solr query filter info" "Clean up search events cache info" "Start New Crawl Job" Click on this API button to see a documentation of the POST request parameter for crawl starts. Expert Crawl Start Start Crawling Job: You can define URLs as start points for Web page crawling and start crawling here. "Crawling" means that YaCy will download the given website, extract all links in it and then download the content behind these links. This is repeated as long as specified under "Crawling Depth". Crawl Job A Crawl Job consist of one or more start point, crawl limitations and document freshness rules. Start Point One Start URL or a list of URLs:<br/>(must start with http:// https:// ftp:// smb:// file://) Define the start-url(s) here. You can submit more than one URL, each line one URL please. Each of these URLs are the root for a crawl start, existing start URLs are always re-loaded. Other already visited URLs are sorted out as "double", if they are not allowed using the re-crawl option. From Link-List of URL From Sitemap From File (enter a path<br/>within your local file system) Index Attributes Add Crawl result to collection<br>(important for Index Pack generation) A crawl result can be tagged with names which are candidates for a collection request. Do not use underline '_' in collection name, use '-' instead. When useful, add a language code to the collection name, e.g. 'top-100-en'. Time Zone Offset The time zone is required when the parser detects a date in the crawled web page. Content can be searched with the on: - modifier which requires also a time zone when a query is made. To normalize all given dates, the date is stored in UTC time zone. To get the right offset from dates without time zones to UTC, this offset must be given here. The offset is given in minutes; Time zone offsets for locations east of UTC must be negative; offsets for zones west of UTC must be positve. Crawler Filter These are limitations on the crawl stacker. The filters will be applied before a web page is loaded. Indexing This enables indexing of the webpages the crawler will download. This should be switched on by default, unless you want to crawl only to fill the Document Cache without indexing. index text index media Do Remote Indexing If checked, the crawler will contact other peers and use them as remote indexers for your crawl. If you need your crawling results locally, you should switch this off. Only senior and principal peers can initiate or receive remote crawls. <strong>A YaCyNews message will be created to inform all peers about a global crawl</strong>, so they can omit starting a crawl with the same start point. Remote crawl results won't be added to the local index as the remote crawler is disabled on this peer. Describe your intention to start this global crawl (optional) This message will appear in the 'Other Peer Crawl Start' table of other peers. Crawling Depth This defines how often the Crawler will follow links (of links..) embedded in websites. 0 means that only the page you enter under "Starting Point" will be added to the index. 2-4 is good for normal indexing. Values over 8 are not useful, since a depth-8 crawl will index approximately 25.600.000.000 pages, maybe this is the whole WWW. also all linked non-parsable documents Unlimited crawl depth for URLs matching with Maximum Pages per Domain You can limit the maximum number of pages that are fetched and indexed from a single domain with this option. You can combine this limitation with the 'Auto-Dom-Filter', so that the limit is applied to all the domains within the given depth. Domains outside the given depth are then sorted-out anyway. Use Page-Count misc. Constraints A questionmark is usually a hint for a dynamic page. URLs pointing to dynamic content should usually not be crawled. However, there are sometimes web pages with static content that is accessed with URLs containing question marks. If you are unsure, do not check this to avoid crawl loops. Following frames is NOT done by Gxxg1e, but we do by default to have a richer content. 'nofollow' in robots metadata can be overridden; this does not affect obeying of the robots.txt which is never ignored. Accept URLs with query-part ('?'): Obey html-robots-noindex: Obey html-robots-nofollow: Media Type detection Not loading URLs with unsupported file extension is faster but less accurate. Indeed, for some web resources the actual Media Type is not consistent with the URL file extension. Here are some examples: Do not load URLs with an unsupported file extension Always cross check file extension against Content-Type header Load Filter on URLs Example: to allow only urls that contain the word 'science', set the must-match filter to '.*science.*'. You can also use an automatic domain-restriction to fully crawl a single domain. must-match Restrict to start domain(s) Restrict to sub-path(s) Use filter (must not be empty) must-not-match Load Filter on URL origin of links Example: to allow loading only links from pages on example.org domain, set the must-match filter to '.*example.org.*'. Load Filter on IPs Must-Match List for Country Codes Crawls can be restricted to specific countries. This uses the country code that can be computed from the IP of the server that hosts the page. The filter is not a regular expressions but a list of country codes, separated by comma. no country code restriction Document Filter These are limitations on index feeder. The filters will be applied after a web page was loaded. Filter on URLs that <b>must not match</b> with the URLs to allow that the content of the url is indexed. No Indexing when Canonical present and Canonical != URL Filter on Content of Document<br/>(all visible text, including camel-case-tokenized url and title) Filter on Document Media Type (aka MIME type) that <b>must match</b> with the document Media Type (also known as MIME Type) to allow the URL to be indexed. Each parsed document is checked against the given Solr query before being added to the index. The embedded local Solr index must be connected to use this kind of filter. Content Filter These are limitations on parts of a document. The filter will be applied after a web page was loaded. You can choose to: Evaluate by default Use all words in document by default until a CSS class as listed below appears; then ignore all Ignore by default Ignore all words in document by default until a CSS class as listed below appears, then evaluate all Filter div or nav class names comma-separated list of &lt;div&gt; or &lt;nav&gt; element class names which should be filtered out/in according to switch above. Clean-Up before Crawl Start Clean up search events cache Check this option to be sure to get fresh search results including newly crawled documents. Beware that it will also interrupt any refreshing/resorting of search results currently requested from browser-side. No Deletion After a crawl was done in the past, document may become stale and eventually they are also deleted on the target host. To remove old files from the search index it is not sufficient to just consider them for re-load but it may be necessary to delete them because they simply do not exist any more. Use this in combination with re-crawl while this time should be longer. Do not delete any document before the crawl is started. Delete sub-path For each host in the start url list, delete all documents (in the given subpath) from that host. Delete only old Treat documents that are loaded ago as stale and delete them before the crawl is started. Double-Check Rules No&nbsp;Doubles A web crawl performs a double-check on all links found in the internet against the internal database. If the same url is found again, then the url is treated as double when you check the 'no doubles' option. A url may be loaded again when it has reached a specific age, to use that check the 're-load' option. Never load any page that is already known. Only the start-url may be loaded again. Re-load ago as stale and load them again. If they are younger, they are ignored. Document Cache Store to Web Cache This option is used by default for proxy prefetch, but is not needed for explicit crawling. Policy for usage of Web Cache The caching policy states when to use the cache during crawling: <b>no&nbsp;cache</b>: never use the cache, all content from fresh internet source; <b>if&nbsp;fresh</b>: use the cache if the cache exists and is fresh using the proxy-fresh rules; <b>if&nbsp;exist</b>: use the cache if the cache exist. Do no check freshness. Otherwise use online source; <b>cache&nbsp;only</b>: never go online, use all content from cache. If no cache exist, treat content as unavailable no&nbsp;cache if&nbsp;fresh if&nbsp;exist cache&nbsp;only Robot Behaviour Use Special User Agent and robot identification Because YaCy can be used as replacement for commercial search appliances (like the Google Search Appliance aka GSA) the user must be able to crawl all web pages that are granted to such commercial platforms. Not having this option would be a strong handicap for professional usage of this software. Therefore you are able to select alternative user agents here which have different crawl timings and also identify itself with another user agent and obey the corresponding robots rule. Enrich Vocabulary Scraping Fields You can use class names to enrich the terms of a vocabulary based on the text content that appears on web pages. Please write the names of classes into the matrix. Vocabulary Class "Scan" Network Scanner YaCy can scan a network segment for available http, ftp and smb server. You must first select a IP range and then, after this range is scanned, it is possible to select servers that had been found for a full-site crawl. Scan the network Scan Range Scan sub-range with given host Do not use intranet scan results, you are not in an intranet environment! All known hosts in the search index (/31 subnet recommended!) Subnet /31 (only the given host(s)) /24 (254 addresses) /20 (4064 addresses) /16 (65024 addresses) Time-Out ms Scan Cache accumulate scan results with access type "granted" into scan cache (do not delete old scan result) Service Type ftp smb http https Scheduler run only a scan scan and add all sites with granted access automatically. This disables the scan cache accumulation. &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;Look every minutes hours days again and add new sites automatically to indexer. Sites that do not appear during a scheduled scan period will be excluded from search results. "empty" "Show all links" "Start New Crawl" Site Crawling Site Crawler: Download all web pages from a given domain or base URL. Site Crawl Start Site Start URL&nbsp;(must start with<br/>http:// https:// ftp:// smb:// file://) Link-List of URL Sitemap URL Path load all files in domain load only files in a sub-path of given url Limitation not more than documents Collection Start Hints Crawl Speed Limitation No more that four pages are loaded from the same host in one second (not more that 120 document per minute) to limit the load on the target server. Target Balancer A second crawl for a different host increases the throughput to a maximum of 240 documents per minute since the crawler balances the load over all hosts. High Speed Crawling A 'shallow crawl' which is not limited to a single host (or site) can extend the pages per minute (ppm) rate to unlimited documents per minute when the number of target hosts is high. Scheduler Steering "API" "Pages Per Minute" "Latency Factor" "Max same Host in queue" "set" "Set PPM to the default minimum value" "Set PPM to the default maximum value" "Terminate" "show link structure" "hide graphic" Click on this API button to see an XML with information about the crawler status Crawler (Please enable JavaScript to automatically update this page!) Queues Queue Size Local Crawler Limit Crawler Remote Crawler No-Load Crawler Terminate All Index Size Database Entries Seg-<br/>ments Citations<br/>(reverse link index) RWIs<br/>(P2P Chunks) Progress Indicator Level Speed / PPM<br/>(Pages Per Minute) <abbr title="Pages Per Minute">PPM</abbr> <abbr title="Latency Factor">LF</abbr> <abbr title="Max same Host in queue">MH</abbr> Crawler PPM Postprocessing Progress pending: Traffic (Crawler) MB Load Error with profile management. Please stop YaCy, delete the file DATA/PLASMADB/crawlProfiles0.db and restart. Application not yet initialized. Sorry. Please wait some seconds and repeat the request. filter. it may take some seconds until the first result appears there.</strong> No embedded local Solr index is connected. This is required to use a Solr query filter. The Solr filter query syntax is not valid : Could not parse the Solr filter query : You asked for remote indexing, but remote crawl results won't be added to the local index as the remote crawler is currently disabled on this peer. Name Count Status Running Crawled Pages "Load" "Deactivate" "Remove" "Activate" Knowledge Loader YaCy can use external libraries to enable or enhance some functions. These libraries are not included in the main release of YaCy because they would increase the application file too much. You can download additional files here. Geolocalization Geolocalization will enable YaCy to present locations from OpenStreetMap according to given search words. GeoNames With this file it is possible to find cities all over the world. Content cities with a population &gt; 1000 all over the world Download from Storage location Status not loaded loaded deactivated Action Result loaded and activated dictionary file deactivated and removed dictionary file deactivated dictionary file activated dictionary file cities with a population &gt; 5000 all over the world cities with a population &gt; 100000 all over the world (the set is is reduced to cities &gt; 100000) OpenGeoDB With this file it is possible to find locations in Germany using the location (city) name, a zip code, a car sign or a telephone pre-dial number. Downloaded from loaded - can be upgraded using the Load button for the new URL loaded and upgraded dictionary file Suggestions Suggestion dictionaries will help YaCy to provide better suggestions during the input of search words DeReWo - Korpusbasierte Grund-/Wortformenlisten (German) of 'Institut f&uuml;r Deutsche Sprache' This file provides 100000 most common german words for suggestions Synonyms Synonyms are used to find not only the searched word but also their synonyms. This is done by adding all synonyms of words in documents to the document and searching the synonyms as well. OpenThesaurus - German Thesaurus from http://www.openthesaurus.de The data from this source was converted to the YaCy synonym file format and part of the YaCy distribution. Deactivated Activated Moby Lexicon - English Thesaurus from https://www.gutenberg.org/ebooks/3202 Russian Thesaurus The data was converted to the YaCy synonym file format and part of the YaCy distribution. YaCy: Tutorial Tutorial You are using the administration interface of your own search engine. You can create your own search index with YaCy. To learn how to do that, watch one of the demonstration videos below: twitter this video More Tutorials "Delete Subpath" "Re-load load-failure docs (404s etc)" "Directory" "Delete Load Errors" Index Browser Host/URL Browse Host Host List URLs Count Colors: Documents without Errors Pending in Crawler Crawler Excludes Load Errors Host Analysis Add to blacklist Path stored linked pending excluded failed Metadata link, detected from context load &amp; index indexed loading Administration Options Delete all from index "Show URL Entries for Word" "Show URL Entries for Word-Hash" "Generate List" "List Selected URLs" "Delete Word" "Transfer to other peer" "Delete reference to selected URLs" "Add selected URLs to blacklist" "Add selected domains to blacklist" Reverse Word Index Administration RWI Retrieval (= search for a single word) Retrieve by Word: Retrieve by Word-Hash: Limitations Index Reference Size No reference size limitation (this may cause strong CPU load when words are searched that appear very often) Limitation of number of references per word: (this causes that old references are deleted if that limit is reached) Set References Limit Search result: total URLs appearance in in link type document type description title creator subject url emphasized image audio video app index of Selection Display URL List Number of lines: all lines Word Deletion delete also the referenced URL (recommended, may produce unresolved references at other word indexes but they do not harm) for every resolvable and deleted URL reference, delete the same reference at every other word where the reference exists (very extensive, but prevents further unresolved references) Transfer RWI to other Peer Transfer by Word-Hash: to Peer: select or enter a hash or peer name: Sequential List of Word-Hashes: No URL entries related to this word hash Resource Negative Ranking Factors Positive Ranking Factors props Reverse Normalized Weighted Ranking Sum hash dom length url comps url length pos in text pos of phrase pos in phrase term frequency authority date words in title words in text local links remote links hitcount unresolved URL Hash Deletion of selected URLs Blacklist Extension "API" "Show Details for URL" "Show Details for URL-Hash" "Delete" "Optimize Solr" "Shut Down and Re-Start Solr" "Generate Statistics" "delete all" "Show Content" "Delete URL" "Delete URL and remove all references from words" Click the API icon to see an example call to the search rss API. URL Database Administration URL Retrieval Retrieve by URL: Retrieve by URL-Hash: Cleanup Index Deletion Delete local search index (embedded Solr and old Metadata) Delete remote solr index Delete RWI Index (DHT transmission words) Delete Citation Index (linking between URLs) Delete First-Seen Date Table Delete HTTP &amp; FTP Cache Stop Crawler and delete Crawl Queues Delete robots.txt Cache Optimize Solr merge to max. segments Reboot Solr Core This feature is available when using exclusively a local embedded Solr. Statistics about top-domains in URL Database Show top domains from all URLs. Domain URLs this may produce unresolved references at other word indexes but they do not harm delete the reference to this url at every other word where the reference exists (very extensive, but prevents unresolved references) Loader Queue The loader set is empty Initiator Depth Status URL "show more" "clear list" Rejected URLs Time URL Fail-Reason "API" "Delete" Click on this API button to see an XML with information about the crawler latency and other statistics. This crawler queue is empty Delete Entries: Initiator Profile Depth Modified Date Anchor Name URL Count Delta/ms Host "Simulate Deletion" "no actual deletion, generates only a deletion count" "Engage Deletion" "simulate a deletion first to calculate the deletion count" "engaged" Index Deletion Deletions are made concurrently which can cause that recently deleted documents are not yet reflected in the document count. Index deletion will not immediately reduce the storage size on disk because entries are only marked as deleted in a first step. Delete by URL Matching Delete all documents within a sub-path of the given urls. That means all documents must start with one of the url stubs as given here. One URL stub, a list of URL stubs<br/>or a regular expression Matching Method sub-path of given URLs matching with regular expression Delete by Age Delete all documents which are older than a given time period. Time Period All documents older than years months days hours Age Identification load date last-modified Delete Collections Delete all documents which are inside specific collections. Not Assigned Delete all documents which are not assigned to any collection Assigned Delete all documents which are assigned to the following collection(s) Delete by Solr Query This is the most generic option: select a set of documents using a solr query. Core "Create Dump" "Restore Dump" Solr Index Export/Import Dump and Restore of Solr Index This feature is available only when a local embedded Solr is active. (This may take several minutes. Please be patient and wait until the page reloads.) Dump File (full path) Could not create the Solr dump : no embedded Solr is available. An error occurred while trying to create the Solr dump. Successfully restored Solr index from dump file! Could not restore the Solr dump : no embedded Solr is available. An error occurred while trying to restore the Solr dump. "Export" Index Export Loaded URL Export Export Path URL Filter query maximum age (seconds) maximum number of records per chunk if exceeded: several chunks are stored; -1 = unlimited (makes only one chunk) Export Size full size, all fields: minified; only fields sku, date, title, description, text_t Export Format Full URL List: Plain Text List (URLs only) HTML (URLs with title) Only Domain: Plain Text List (domains only) HTML (domains as URLs, no title) Only Text: Fulltext of Search Index Text Import this file by moving it to DATA/PACKS/load "Set" Index Sources &amp; Targets YaCy supports multiple index storage locations. As an internal indexing database a deep-embedded multi-core Solr is used and it is possible to attach also a remote Solr. Solr Search Index Lazy Value Initialization If checked, only non-zero values and non-empty strings are written to Solr fields. Use deep-embedded local Solr This will write the YaCy-embedded Solr index which is stored within the YaCy DATA directory. The Solr native search interface is accessible at /solr/select?q=*:*&amp;start=0&amp;rows=3&amp;core=collection1 for the default search index (core: collection1) and at If you switch off this index, a remote Solr must be activated. Use remote Solr server(s) This external Solr can be used instead of the internal Solr. It can also be used additionally to the internal Solr, then both Solr indexes are mirrored. Allow self-signed certificates Tick this when the remote Solr server is password protected and is requested over HTTPS but provides only a self-signed certificate (not a validated one by an official Certificate Authority). The Solr URL could be for example something like <i>https://user:password@localhost:8984/solr</i>. Solr Hosts Solr Host Administration Interface Index Size Solr URL(s) You can set one or more Solr targets here which are accessed as a shard. For several targets, list them using a ',' (comma) as separator. The set of remote targets are used as shards of a complete index. The host part of the url is used as key for a hash function which selects one of the shards (one of your remote servers). When a search request is made, all servers are accessed synchronously and the result is combined. Sharding Method write-enabled (if unchecked, the remote server(s) will only be used as search peers) Web Structure Index The web structure index is used for host browsing (to discover the internal file/folder structure), ranking (counting the number of references) and file search (there are about forty times more links from loaded pages than in documents of the main search index). use citation reference index (lightweight and fast) use webgraph search index (rich information in second Solr core) Peer-to-Peer Operation The 'RWI' (Reverse Word Index) is necessary for index transmission in distributed mode. For portal or intranet mode this must be switched off. support peer-to-peer index transmission (DHT RWI index) Block known error URLs in DHT Reject URLs/RWIs with known errors from peers. Disable to opt out. Retry after (days) for temporary errors; permanent errors stay blocked. Permanent error statuses comma-separated (default: 404,410,-1; -1=DNS/network errors) "Import JsonList File" "Stop" JSON List Index Dump File Import No import thread is running, you can start a new thread here JsonList File Selection: select an jsonlist file (which may be gz compressed) File: or Url: Import Process Thread: JsonList File: Processed: Speed: Running Time: Remaining Time: "Uniform Resource Locator" "Dump file path on this YaCy server file system, or any remote URL" "Import MediaWiki Dump" MediaWiki Dump Import No import thread is running, you can start a new thread here Error : dump <abbr title="Uniform Resource Locator">URL</abbr> is malformed. MediaWiki Dump File Selection Dumps can be stored in the local file system or on a remote server in XML format and may be compressed in gz or bz2. Dump file path or <abbr title="Uniform Resource Locator">URL</abbr> Import only when modified since last import When checked, the dump file is imported only if its last modified date is unknown or is after the last import execution date on this same file When the import is started, the following happens: The dump is extracted on the fly and wiki entries are translated into Dublin Core data format. The output looks like this: Each 10000 wiki records are combined in one output file which is written to /DATA/PACKS/load into a temporary file. When each of the generated output file is finished, it is renamed to a .xml file Each time a xml pack file appears in /DATA/PACKS/load, the YaCy indexer fetches the file and indexes the record entries. When a pack file is finished with indexing, it is moved to /DATA/PACKS/loaded You can recycle processed pack files by moving them from /DATA/PACKS/loaded to /DATA/PACKS/load Import Process Thread: started running Dump: Processed: Speed: Running Time: Remaining Time: "Load Selected Sources" Source Import List Thread Processed<br />Chunks Imported<br />Records Complete at<br /># Records Speed<br />(records/second) "Import OAI-PMH source" "import this source" "import from a list" OAI-PMH Import Single request import This will submit only a single request as given here to a OAI-PMH server and imports records into the index Source: Processed: ResumptionToken: Import all Records from a server Import all records that follow according to resumption elements into index or Import started! "Import Warc File" "Stop" Web Archive File Import No import thread is running, you can start a new thread here Warc File Selection: select an warc file (which may be gz compressed) You can download warc archives for example here File: or Url: Collection: Import Process Thread: Warc File: Processed: Speed: Running Time: Remaining Time: "Import ZIM File" "Stop" ZIM File Import No import thread is running, you can start a new thread here Zim File Selection: select a '.zim' file You can download ZIM files for example here File: Collection: Import Process Thread: ZIM File: Processed: Speed: Running Time: Remaining Time: YaCy Pack Downloader Available Packs Source Repo ID File Process "info" "Generate Data Pack" YaCy Pack Generator Index Pack Generator Set a Category (this goes into the filename) mix - a mix of document types, for content from wide web crawls core - technical documentation, operating systems, computer hardware, open source and free software, manuals, protocol standards scroll - non-technical documents: knowledge, encyclopedia, linguistic corpora, dictionaries, translation memories, texts, non-fiction books, historical books regula - non-technical standards: industry standards, laws, rules, compliance gem - research, papers, university publications, science fiction - fictional documents: movies, stories, series, books (fiction, science-fiction) map - geological data, geolocation-data, earth/world information echo – micro-content (tweets, toots, short headlines, SMS corpora), podcasts, radio archives, audio lectures, spoken-word datasets, logs, incidents, telemetry spirit – related to non-textual data (possibly only metadata): art, music, game assets, creative-commons media (non-text culture loot) vault - sensitive data: secrets, leaks, non-public documents, security advisories Index Collection the collection name is used as part of the filename to describe the content. Exception: if the collection is "user", then you can name the content with a slug. Slug - describe the content<br>(only if collection is "user") This will become a part of the filename, spaces will be replaced by "-"; must not be empty; should end with a language description, e.g. "-en" URL Filter Search Query - Export Format This JSON is an elasticsearch index dump format and can be bulk-imported to elasticsearch. Here is an example for opensearch, using docker: Start docker container of opensearch: Unblock index creation: Create the search index: Bulk-upload the index file: Make a search, get 10 results, search in fields text_t, title, description with boosts: JSON (Rich and full-text Elasticsearch data, one document per line in one flat JSON file) XML (Rich and full-text Solr data, one document per line in one large xml file, can be processed with shell tools, can be imported with DATA/PACKS/load/) XML (RSS) Import this file by moving it to DATA/PACKS/load Pack List Pack Process Size (KB) YaCy Pack Manager Pack Folders Packs: Hold List Size (KB) Process Packs: Load List Packs: Loaded List "refresh page" "start reindex job now" "stop reindexing" "Simulate" "Check only how many documents would be selected for recrawl" "Set defaults" "Reset to default values" "start recrawl job now" "update" "stop recrawl job" "Automatically refreshing" "An error occurred while trying to refresh automatically" "URLs added to the crawler queue for recrawl" "URLs rejected for some reason by the crawl stacker or the crawler queue. Please check the logs for more details." Field Re-Indexing In case that an index schema of the embedded/local index has changed, all documents with missing field entries can be indexed again with a reindex job. Documents in current queue Documents processed current select query Remaining field list reindex documents containing these fields: Field count Re-Crawl Index Documents Searches the local index and selects documents to add to the crawler (recrawl the document). This runs transparent as background job. Documents are added to the crawler only if no other crawls are active and are added in small chunks. Re-crawl works only with an embedded local Solr index! Solr query document(s) selected for recrawl. An error occurred when trying to run the selection query. The Solr index is not connected. Please restart your peer. Include failed URLs Delete URLs to re-crawl documents selected with the given query. Re-Crawl Query Details Documents to process Current Query Edit Solr Query Include failed urls Delete urls Last Re-Crawl job report The job terminated early due to an error when requesting the Solr index. Status Running Shutdown in progress Terminated Query Start time End time Recrawled URLs Rejected URLs Malformed URLs Refresh "API" "active" "disabled" "Required for proper operation" "Set" "reset selection to default" "reindex Solr" The solr schema can also be retrieved as xml here. Click the API icon to see the xml. Just copy this xml to solr/conf/schema.xml to configure solr. Solr Schema Editor If you use a custom Solr schema you may enter a different field name in the column 'Custom Solr Field Name' of the YaCy default attribute name Select a core: Active Attribute Custom Solr Field Name Comment show active show all available show disabled Reindex documents If you unselected some fields, old documents in the index still contain the unselected fields. To physically remove them from the index you need to reindex the documents. Here you can reindex all documents with inactive fields. "Set" Index Sharing Index: distribute&nbsp; receive receive grant default: for each remote peer links/minute&nbsp; words/minute "info" LLM Selection Here you can pick models from an LLM model service to select them as production model. In the "Production Models Matrix" you can then assign each selected model a function inside YaCy Service Selection service Ollama LMStudio OpenAI Open Router This makes a preset to the Hoststub value hoststub you can probably leave this to the default value api_key (not required for Ollama or LMStudio) Services <b>num_ctx</b> is the context window (in tokens) of the inference service &mdash; a per-service value, shared by all models on that endpoint. It is the total budget for prompt <i>plus</i> generated output; YaCy uses it to size prompts so they leave room to generate. The row for the service selected above appears here automatically with its stored (or default) window. This value is <b>advisory</b>: set it to match the window your backend actually serves Context Length setting). YaCy does not enforce it on the backend. num_ctx Model Downloads Production Models Matrix model max_tokens search-answers This model creates answers for search requests chat This model is used in the chat interface and as default for the RAG proxy translation This model can be used to make translations of the web UI classification This model is used to classify prompts to find out what they demand search-query This model produces search queries to YaCy search from prompts in RAG or chat qa-pairs This model can be used to produce query-answer pairs which enhance search from chat prompts tldr-shortener This model is used to make summaries from web content log-report This model evaluates YaCy runtime logs and creates self-enhancement reports thinking we detect thinking only to be able to suppress thinking. thinking is not used in YaCy tooling tooling is required for agentic abilities. vision this enables image recognition in the chat format this is required for classification Actions "Get content of Wiki: crawl wiki pages" Integration in MediaWiki It is possible to insert wiki pages into the YaCy index using a web crawl on that pages. This guide helps you to crawl your wiki and to insert a search window in your wiki pages. Retrieval of Wiki Pages The following form is a simplified crawl start that uses the proper values for a wiki crawl. Just insert the front page URL of your wiki. After you started the crawl you may want to get back to this page to read the integration hints below. <b>URL of the wiki main page</b><br />This is a crawl start point Inserting a Search Window to MediaWiki To integrate a search window into a MediaWiki, you must insert some code into the wiki template. There are several templates that can be used for MediaWiki, but in this guide we consider that you are using the default template, 'MonoBook.php': open skins/MonoBook.php find the line where the default search window is displayed, there are the following statements: Remove that code or set it in comments using '&lt;!--' and '--&gt;' Insert the following code: Check all appearances of static IPs given in the code snippet and replace it with your own IP, or your host name You may want to change the default text elements in the code snippet To see all options for the search widget, look at the more generic description of search widgets at "Get content of phpBB3: crawl forum pages" Integration in phpBB3 It is possible to insert forum pages into the YaCy index using a database import of forum postings. This guide helps you to insert a search window in your phpBB3 pages. Retrieval of phpBB3 Forum Pages using a database export Forum posting contain rich information about the topic, the time, the subject and the author. This information is in an bad annotated form in web pages delivered by the forum software. It is much better to retrieve the forum postings directly from the database. This will cause that YaCy is able to offer nice navigation features after searches. Retrieval of phpBB3 Forum Pages using a web crawl The following form is a simplified crawl start that uses the proper values for a phpbb3 forum crawl. Just insert the front page URL of your forum. After you started the crawl you may want to get back to this page to read the integration hints below. <b>URL of the phpBB3 forum main page</b><br />This is a crawl start point Inserting a Search Window to phpBB3 To integrate a search window into phpBB3, you must insert some code into a forum template. There are several templates that can be used for phpBB3, but in this guide we consider that you are using the default template, 'prosilver': open styles/prosilver/template/overall_header.html Insert the following code right behind the div tag: Check all appearances of static IPs given in the code snippet and replace it with your own IP, or your host name You may want to change the default text elements in the code snippet To see all options for the search widget, look at the more generic description of search widgets at "Show RSS Items" "Add All Items to Index (full content of url)" "Remove Selected Feeds from Scheduler" "Remove All Feeds from Scheduler" "Remove Selected Feeds from Feed List" "Remove All Feeds from Feed List" "Add Selected Feeds to Scheduler" "Add Selected Items to Index (full content of url)" Loading of RSS Feeds RSS feeds can be loaded into the YaCy search index. This does not load the rss file as such into the index but all the messages inside the RSS feeds as individual documents. URL of the RSS feed Preview Indexing Available after successful loading of rss feed in preview once load this feed once now scheduled repeat the feed loading every minutes hours days automatically. collection List of Scheduled RSS Feed Load Targets Title URL/Referrer Recording Last Load Next Load Last Count All Count Avg. Update/Day Available RSS Feed List Author Description Language Date Time-to-live Docs State URL new enqueued indexed Attached media "delete this report" Log Reports run report now Generating report from the current-hour log lines &mdash; the LLM call can take a while &hellip; seconds elapsed No log lines were found for the current hour. No production model is configured for the log-report role. Assign one in the No production model is configured for the log-report role. Log report generation stays inactive until a model is assigned in the Feeds: JSON RSS The report directory does not exist yet. Reports will appear here after the scheduler has generated the first completed hourly report. &times; Report generation in progress &hellip; the report below is completed live while the model is writing No generated log reports were found. "Enter" "Preview" Send message The peer does not respond. It was now removed from the peer-list. Your Message Subject: Text: The peer is alive but cannot respond. Sorry. Preview message The message has not been sent yet! Message: Your message has been sent. The target peer responded: The target peer is alive but did not receive your message. Sorry. Here is a copy of your message, so you can copy it to save it for further attempts: "RSS" "Compose" Messages Compose Message Send message to peer Date From To Subject Action view reply delete From: To: Date: Subject: Message: Action: inbox "API" "Search" "https supported" "Type: Junior | Contact: passive" "Junior passive" "Type: Junior | Contact: direct" "Junior direct" "Type: Junior | Contact: offline" "Junior offline" "Type: Senior | Contact: passive" "senior passive" "Type: Senior | Contact: direct" "Senior direct" "Type: Senior | Contact: offline" "Senior offline" "Type: Principal | Contact: passive | Seed download: possible" "Principal passive" "Type: Principal | Contact: direct | Seed download: possible" "Principal active" "Type: Principal | Contact: offline | Seed download: ?" "Principal offline" "Accept Crawl: no" "no crawl" "Accept Crawl: yes" "crawl possible" "no DHT receive" "DHT Receive: yes" "DHT receive enabled" "Profile updated" "Wiki updated" "Blog updated" "Crawl" "The YaCy Network" "Type: Virgin" "Virgin" "Type: Junior" "Junior" "Type: Senior" "Senior" "Type: Principal" "Principal" "Crawl enabled" "DHT Receive: no" "DHT Receive enabled" "add Peer" "contact current peer from this peer" YaCy Network Network Overview Active&nbsp;Principal&nbsp;and&nbsp;Senior&nbsp;Peers Passive&nbsp;Senior&nbsp;Peers Junior&nbsp;(fragment)&nbsp;Peers Network History The information that is presented on this page can also be retrieved as XML. Click the API icon to see the XML. Manually contacting Peer Search for a peername (RegExp allowed) Hash Name Info Release Age con/h<br/> PPM QPH Last<br/>Seen <strong>UTC</strong><br/>Offset Uptime Links RWIs URLs<br/>for<br/>Remote<br/>Crawl Sent DHT<br/>Word Chunks Sent<br/>URLs Received DHT<br/>Word Chunks Received<br/>URLs Location user agent<br/> send&nbsp;<strong>M</strong>essage/<br/>show&nbsp;<strong>P</strong>rofile/<br/>edit&nbsp;<strong>W</strong>iki/<br/>browse&nbsp;<strong>B</strong>log Network Online Peers Number of<br/>Documents Indexing Speed:<br/>Pages Per Minute (PPM) Query Frequency:<br/>Queries Per Hour (QPH) Last Hour Today Last&nbsp;Week Last&nbsp;Month Now Active Senior Passive Senior Junior (fragment) This Peer Your Peer: Version UTC URLs for<br/>Remote Crawl Sent<br/>DHT Word Chunks Received<br/>DHT Word Chunks Known<br/>Seeds Connects<br/>per hour Indexing<br/>PPM QPH<br/>(public&nbsp;local) QPH<br/>(remote) dark green font senior/principal peers light green font passive peers pink font junior peers red point this peer grey waves crawling activity green radiation strong query activity red lines DHT-out green lines DHT-in Peer Hash Peer IP Peer Port Contacting current peer from another: ip:port <b>Count of Connected Senior Peers</b> in the last two days, scale = 1h <b>Count of all Active Peers Per Day</b> in the last week, scale = 1d <b>Count of all Active Peers Per Week</b> in the last 30d, scale = 7d <b>Count of all Active Peers Per Month</b> in the last 365d, scale = 30d "Incoming News" "Processed News" "Outgoing News" "Published News" Overview Incoming&nbsp;News Processed&nbsp;News Outgoing&nbsp;News Published&nbsp;News This is the YaCyNews system (currently under testing). The news service is controlled by several entry points: A crawl start with activated remote indexing will automatically create a news entry. Other peers may use this information to prevent double-crawls from the same start point. A table with recently started crawls is presented on the Index Create - page A change in the personal profile will create a news entry. You can see recently made changes of profile entries on the Network page, where that profile change is visualized with a '*' beside the 'P' (profile) - selector. Publishing of added or modified translation for the user interface. Other peers may include it in their local translation list. More news services will follow. Above you can see four menus: Only these news will be used to display specific news services as explained above. You can process these news with a button on the page to remove their appearance from the IndexCreate and Network page you can stop the broadcast if you want. Originator Created Category Received Distributed Attributes Performance of Concurrent Processes serverProcessor Objects Thread Queue Size<br />Current Queue Size<br />Maximum Executors:<br />Current Number of Threads Concurrency:<br />Maximum Number of Threads Children Average<br />Block Time<br />Reading Average<br />Exec Time Average<br />Block Time<br />Writing Total<br />Cycles Full Description "PerformanceGraph" Performance Settings for Memory refresh graph simulate short memory status use Standard Memory Strategy Memory Usage Type After Startup After Initializations<br />before GC After Initializations<br />after GC Now before GC after GC Description Max maximum memory that the JVM will attempt to use Available total available memory including free for the JVM within maximum Total total memory taken from the OS Free free memory in the JVM within total amount Used used memory in the JVM within total amount Table RAM Index Table Size Key Value Chunk Size Used Memory Object Index Caches Needed Memory Other Caching Structures Hit Miss Insert Delete DNSCache/Hit (ARC) DNSCache/Miss DNSNoCache HashBlacklistedCache Search Event Cache "Submit New Delay Values" "Re-set to default" "When the system load average is over the specified value, that type of remote search request is not used to fill search results." "Reverse Word Index" "Submit New Values" "Enter New Cache Size" "Enter new Threadpool Configuration" "Total maximum number of simultaneously open connections in the pool" "Number of connections currently being used to execute requests." "Number of reusable idle connections" "Number of connection requests being blocked awaiting a free connection" Performance Settings of Queues and Processes Scheduled tasks overview and waiting time settings: Thread Queue Size Total<br />Block Time Total<br />Sleep Time Total<br />Exec Time Total<br />Cycles Idle<br />Cycles Busy<br />Cycles Short Mem<br />Cycles High CPU<br />Cycles Sleep Time<br />per Cycle<br />(millis) Exec Time<br />per Busy-Cycle<br />(millis) Memory Use<br />per Busy-Cycle<br />(kbytes) Delay between<br />idle loops Delay between<br />busy loops Minimum of<br />Required Memory Maximum of<br />System-Load Full Description milliseconds kbytes load Changes take effect immediately Remote search requests: Type Maximum system load <abbr title="Reverse Word Index">RWI</abbr> Search requests performed on remote peers distributed Reverse Word Index Solr Search requests performed on remote peers Solr indexes Cache Settings: RAM Cache Description Words in RAM cache:<br />(Size in KBytes) This is the current size of the word caches. The indexing cache speeds up the indexing process, the DHT cache holds indexes temporary for approval. The maximum of this caches can be set below. Maximum URLs currently assigned<br />to one cached word: This is the maximum size of URLs assigned to a single word cache entry. If this is a big number, it shows that the caching works efficiently. Maximum age of a word: This is the maximum age of a word in an index in minutes. Minimum age of a word: This is the minimum age of a word in an index in minutes. Maximum number of words in cache: This is is the number of word indexes that shall be held in the ram cache during indexing. When YaCy is shut down, this cache must be flushed to disc; this may last some minutes. Thread Pool Settings: Thread Pool maximum Active current Active Outgoing connections pools settings : Connection Pool Total maximum Current statistics Active Idle Pending General Remote Solr servers "Search event picture" Search Sequence Timing Timing results of latest search request: Query Event Comment Time Delta (ms) Duration (ms) Result-Count The network picture below shows how the latest search query was solved by asking corresponding peers in the DHT: red -&gt; request list alive green -&gt; request has terminated grey -&gt; the search target hash order position(s) (more targets if a dht partition is used) "PerformanceGraph" "Java Virtual Machine" "Set" "Restart now" "Amount of space (in Mebibytes) that should be kept free as steady state" "Mebibyte" "Amount of space (in Megabytes) that should at least be kept free as hard limit" "Distributed Hash Table" "Free space disk autoregulation info" "Maximum amount of space (in Mebibytes) that should be used as steady state" "Maximum amount of space (in Mebibytes) that should be used as hard limit" "Used space disk autoregulation info" "Random Access Memory" "Proper state info" "Exhausted state info" "Reset state" "Manually reset to 'proper' state" "Amount of memory (in Mebibytes) that should at least be free for proper operation" "Save" "Enter New Parameters" Performance Settings refresh graph Memory Settings Memory reserved for <abbr title="Java Virtual Machine">JVM</abbr> MByte Accepted change. This will take effect after <strong>restart</strong> of YaCy. Restart now Resource Observer Free space disk Steady-state minimum <abbr title="Mebibyte">MiB</abbr>. Disable crawls when free space is below. Absolute minimum <abbr title="Mebibyte">MiB</abbr>. Disable <abbr title="Distributed Hash Table">DHT</abbr>-in when free space is below. Autoregulate when absolute minimum limit has been reached. The autoregulation task performs the following sequence of operations, stopping once free space disk is over the steady-state value : delete old releases delete logs delete robots.txt table delete news clear HTCACHE clear citations throw away large crawl queues cut away too large RWIs Used space disk Steady-state maximum <abbr title="Mebibyte">MiB</abbr>. Disable crawls when used space is over. Absolute maximum <abbr title="Mebibyte">MiB</abbr>. Disable <abbr title="Distributed Hash Table">DHT</abbr>-in when used space is over. when absolute maximum limit has been reached. The autoregulation task performs the following sequence of operations, stopping once used space disk is below the steady-state value: <abbr title="Random Access Memory">RAM</abbr> Memory state : proper Enough memory is available for proper operation. <strong aria-describedby="exhaustedStateInfo">exhausted</strong> Within the last eleven minutes, at least four operations have tried to request memory that would have reduced free space within the minimum required. Minimum required <abbr title="Mebibyte">MiB</abbr> free space. Disable <abbr title="Distributed Hash Table">DHT</abbr>-in below. Online Caution Settings: This is the time that the crawler idles when the proxy is accessed, or a local or remote search is done. The delay is extended by this time each time the proxy is accessed afterwards. This shall improve performance of the affected process (proxy or search). seconds since last proxy/local-search/remote-search access.) Online Caution Case indexer delay (milliseconds) after case occurrence Proxy: Local Search: Remote Search: Changes take effect immediately "Set proxy profile" Indexing with Proxy YaCy can be used to 'scrape' content from pages that pass the integrated caching HTTP proxy. When scraping proxy pages then <strong>no personal or protected page is indexed</strong>; those pages are detected by properties in the HTTP header (like Cookie-Use, or HTTP Authorization) or by POST-Parameters (either in URL or as HTTP protocol) and automatically excluded from indexing. Proxy Auto Config: this controls the proxy auto configuration script for browsers at http://localhost:8090/autoconfig.pac whether the proxy should only be used for .yacy-Domains Proxy pre-fetch setting: this is an automated html page loading procedure that takes actual proxy-requested URLs as crawling start points for crawling. Prefetch Depth A prefetch of 0 means no prefetch; a prefetch of 1 means to prefetch all embedded URLs, but since embedded image links are loaded by the browser this means that only embedded href-anchors are prefetched additionally. Store to Cache It is almost always recommended to set this on. The only exception is that you have another caching proxy running as secondary proxy and YaCy is configured to used that proxy in proxy-proxy - mode. Do Local Text-Indexing If this is on, all pages (except private content) that passes the proxy is indexed. Do Local Media-Indexing This is the same as for Local Text-Indexing, but switches only the indexing of media content on. Do Remote Indexing If checked, the crawler will contact other peers and use them as remote indexers for your crawl. If you need your crawling results locally, you should switch this off. Only senior and principal peers can initiate or receive remote crawls. Please note that this setting only take effect for a prefetch depth greater than 0. Proxy generally Path The path where the pages are stored (max. length 300) Size The size in MB of the cache. <strong>The file DATA/PLASMADB/crawlProfiles0.db is missing or corrupted. Please delete that file and restart.</strong> <strong>Caching is now off on <strong>Local Text Indexing is now <strong>Local Media Indexing is now <strong>Remote Indexing is now Changes will take effect after restart only. You can see a snapshot of recently indexed pages Quickly adding Bookmarks: Simply drag and drop the link shown below to your Browsers Toolbar/Link-Bar. If you click on it while browsing, the currently viewed website will be inserted into the YaCy crawling queue for indexing. Crawl with YaCy Title: Link: Status: URL successfully added to Crawler Queue Malformed URL Wire RAG Retrieval Tune how YaCy constructs prompts and search queries for Retrieval Augmented Generation. System Prompt This is sent as the system message for chats. Keep it concise and friendly. User Retrieval Prefix Prepended before attached search snippets in RAG mode to tell the LLM how to use them. Query Generator Prefix Prompt given to the model that generates search queries from user requests. Search Document Max Length Maximum character length of the virtual search document used as RAG attachment and as the `search` tool result. Content beyond this limit is cut off. Default: 30000. Save RAG Settings "info" "Set as Default Ranking" "Re-Set to Built-In Ranking" RWI Ranking Configuration The document ranking influences the order of the search result entities. A ranking is computed using a number of attributes from the documents that match with the search word. The attributes are first normalized over all search results and then the normalized attribute is multiplied with the ranking coefficient computed from this list. The ranking coefficient grows exponentially with the ranking levels given in the following table. If you increase a single value by one, then the strength of the parameter doubles. Pre-Ranking There are two ranking stages: first all results are ranked using the pre-ranking and from the resulting list the documents are ranked again with a post-ranking. The two stages are separated because they need statistical information from the result of the pre-ranking. Post-Ranking "Set Boost Function" "Re-Set to default" "Set Boost Query" "Set Filter Query" "Set Field Boosts" Solr Ranking Configuration These are ranking attributes for Solr. This ranking applies for internal and remote (P2P or shard) Solr access. Select a profile: Boost Function A Boost Function can combine numeric values from the result document to produce a number which is multiplied with the score value from the query result. Example: to order by date, use "recip(ms(NOW,last_modified),3.16e-11,1,1)", to order by crawldepth, use "div(100,add(crawldepth_i,1))". Boost Query The Boost Query is attached to every query. Use this to statically boost specific content in the index. Example: "fuzzy_signature_unique_b:true^100000.0f" means that documents, identified as 'double' are ranked very bad and appended to the end of all results (because the unique are ranked high). Filter Query The Filter Query is attached to every query. Use this to statically add a selection criteria to reduce the set of results. Example: "http_unique_b:true AND www_unique_b:true" will filter out all results where urls appear also with/without http(s) and/or with/without 'www.' prefix. Solr Boosts field not in local index (boost has no effect) Regex Test Test String Regular Expression Result no match match "Save" Remote Crawler The remote crawler is a process that requests urls from other peers. Peers offer remote-crawl urls if the flag 'Do Remote Indexing' is switched on when a crawl is started. Remote Crawler Configuration Your peer cannot accept remote crawls because you need senior or principal peer status for that! Accept Remote Crawl Requests Perform web indexing upon request of another peer. Load with a maximum of pages per minute Peers offering remote crawl URLs If the remote crawl option is switched on, then this peer will load URLs from the following remote peers: Name URLs for<br/>Remote<br/>Crawl Release PPM QPH Last<br/>Seen <strong>UTC</strong><br/>Offset Uptime Links RWIs Age "Submit" "Set defaults" "Reset to defaults settings" limitations Local Search access rate limitations You can configure here limitations on access rate to this peer search interface by unauthenticated users and users without extended search right YaCy search Access rate limitations to this peer search interface. When a user with limited rights (unauthenticated or without extended search right) exceeds a limit, the search is blocked. Max searches in 3s Max searches in 1mn Max searches in 10mn Peer-to-peer search Access rate limitations to the peer-to-peer search mode. When a user with limited rights (unauthenticated or without extended search right) exceeds a limit, the search scope falls back to only this local peer index. Max searches in 10mn Peer-to-peer search with JavaScript results resorting Access rate limitations to the peer-to-peer search mode with browser-side JavaScript results resorting enabled When a user with limited rights (unauthenticated or without extended search right) exceeds a limit, results resorting becomes only applicable on demand, server-side. Remote snippet load Limitations on snippet loading from remote websites. When a user with limited rights (unauthenticated or without extended search right) exceeds a limit, the snippets fetch strategy falls back to 'CACHEONLY' Max searches in 3s <em id="changeInfo">Changes will take effect immediately.</em> "Add Selected Servers to Crawler" Network Scanner Monitor The following servers can be searched: Available server within the given IP range Protocol IP URL Access Process inaccessible empty granted denied not in index indexed Settings Receipt: No information has been submitted Nothing changed. Error with submitted information. The user name must be given. Your request cannot be processed.<br />Nothing changed. The password redundancy check failed. You have probably mistyped your password. <strong>Shutting down.</strong><br />Application will terminate after working off all crawling tasks. Your administration account setting has been made. <strong>Your proxy access setting has been changed. Your proxy account check has been disabled.</strong> The new proxy IP filter is set to The proxy port is: Port rebinding will be done in a few seconds. Your proxy access setting has been changed. If you open any public web page through the proxy, you must log-in. Port rebinding will be done in a view seconds. Auto pop-up of the Status page is now <strong>disabled</strong> Auto pop-up of the Status page is now <strong>enabled</strong> The Peer Name is: Your static Ip(or DynDns) is: Your public port is: <strong>Seed Settings changed. You are now a principal peer. Seed Settings changed, but something is wrong. Seed Uploading was deactivated automatically. Please return to the settings page and modify the data. The remote-proxy setting has been changed The new setting is effective immediately, you don't need to re-start. <strong>The submitted peer name is already used by another peer. Please choose a different name.</strong> The Peer name has not been changed. Your Peer Language is: <strong>The submitted peer name is not well-formed. Please choose a different name.</strong> The Peer name has not been changed. Peer names must not contain characters other than (a-z, A-Z, 0-9, '-', '_') and must not be longer than 80 characters. Seed Upload method was changed successfully. Seed Upload Method: Seed File URL: Your proxy networking settings have been changed. Transparent Proxy Support is: Always Fresh is: Send via header is: Send X-Forwarded-For header is: Your message forwarding settings have been changed. Message Forwarding Support is: Message Forwarding Command: Recipient Address: Invalid IP-Number filter: Your crawler settings have been changed. Generic Settings: Crawler timeout: http Crawler Settings: Maximum HTTP Filesize: ftp Crawler Settings: Maximum FTP Filesize: smb Crawler Settings: Maximum SMB Filesize: Maximum file Filesize: Invalid crawler timeout value: Invalid maximum file size for http crawler: Invalid maximum file size for ftp crawler: HTTPS port is now: the change will take effect after restart. URL Proxy settings have been saved. Debug/Analysis settings have been saved. Referrer policy settings have been saved. The ports are now configured as follows (active on next start). HTTP port HTTPS port Shutdown port Compression settings have been saved. HTTP client settings have been saved. Your need to restart YaCy to activate the changes. "Submit" Crawler Settings <strong>Generic Crawler Settings</strong>: Timeout: HTTP Crawler Settings: Maximum Filesize: Please note that if the crawler uses content compression, this limit is used to check the compressed content size.</em> <strong>FTP Crawler Settings</strong>: <strong>SMB Crawler Settings</strong>: <strong>Local File Crawler Settings</strong>: Changes will take effect immediately. "Extensible Markup Language" "Distributed Hash Table" "Reverse Word Index" "Submit" Debug/Analysis Settings Be careful with these advanced settings, they can deeply affect the search process! You probably don't need to modify them for normal use. Solr communication Enable remote Solr binary responses When checked (default), responses from remote Solr index instances are transferred using an efficient binary data format. When unchecked, responses are transferred as <abbr title="Extensible Markup Language">XML</abbr>, which can be captured and parsed by any external XML aware tool for debug/analysis. Search data sources By default all data sources are enabled to obtain search results, but you can here disable one or more ones to check the behavior of the process. Local <abbr title="Distributed Hash Table">DHT</abbr>/<abbr title="Reverse Word Index">RWI</abbr> Local Solr index Remote <abbr title="Distributed Hash Table">DHT</abbr>/<abbr title="Reverse Word Index">RWI</abbr> Remote Solr indexes Search testing tweaks Override <abbr title="Distributed Hash Table">DHT</abbr> peers selection by local only When checked, the remote <abbr title="Distributed Hash Table">DHT</abbr> peers selection is overridden and only the local peer is selected to provide remote DHT search results. Override Solr peers selection by local only When checked, the remote Solr peers selection is overridden and only this peer is selected to provide remote Solr search results. Ranking information Show search results scores When checked, the raw ranking score value is displayed for each text search result in the HTML results page. Text snippets statistics Enable text snippets statistics <em id="submitInfo">Changes will take effect immediately.</em> "Transport Layer Security" "Server Name Indication" "Submit" HTTP client settings You can configure here some advanced settings of the clients used by YaCy to handle outgoing HTTP connections. About Server Name Indication (SNI): this extension to the <abbr title="Transport Layer Security">TLS</abbr> protocol must be enabled to load some https URLs (for websites deployed with different certificates and host names on the same shared IP address), otherwise loading fails with errors such as Received fatal alert: handshake_failure But it can be necessary to disable it in order to load some https URLs served by old and misconfigured web servers, otherwise loading fails with the exception javax.net.ssl.SSLProtocolException: "handshake alert: unrecognized_name" Controlling <abbr title="Server Name Indication">SNI</abbr> extension activation can also be done with the JVM option jsse.enableSNIExtension , but in that case a server restart is required when you want to modify the setting and it is not customizable per http client (general or for remote Solr). General HTTP client Configuration settings for the main HTTP client, used notably to crawl websites and communicate with other YaCy peers. Enable <abbr title="Server Name Indication">SNI</abbr> extension to <abbr title="Transport Layer Security">TLS</abbr> Remote Solr HTTP client Configuration settings for the specific HTTP client dedicated to communications with remote Solr servers (located on other YaCy peers or eventually owned by this one when it is configured to use a remote Solr index). <em id="submitInfo">Changes will take effect immediately.</em> "Submit" Message Forwarding With this settings you can activate or deactivate forwarding of yacy-messages via email. Enable message forwarding Enabling/Disabling message forwarding via email. Forwarding Command <i>The command-line program that should be used to forward the message. e.g.:</i> Forwarding To <i>The recipient email-address. Changes will take effect immediately. "Submit" Remote Proxy (optional) YaCy can use another proxy to connect to the internet. You can enter the address for the remote proxy here: Use remote proxy Enables the usage of the remote proxy by yacy Use remote proxy for HTTPS Specifies if YaCy should forward ssl connections to the remote proxy. Remote proxy host The ip address or domain name of the remote proxy Remote proxy port the port of the remote proxy Remote proxy user Remote proxy password No-proxy addresses IP addresses for which the remote proxy should not be used Changes will take effect immediately. "Submit" "change" Proxy Settings Transparent Proxy With this you can specify if YaCy can be used as transparent proxy. <em>Hint: On linux you can configure your firewall to transparently redirect all http traffic through yacy using this iptables rule</em>: Always Fresh If unchecked, the proxy will act using Cache Fresh / Cache Stale rules. If checked, the cache is always fresh which means that a page is never loaded again if it was already stored in the cache. However, if the page does not exist in the cache, it will be loaded in any case. Send "Via" Header http header according to RFC 2616 Sect 14.45. Send "X-Forwarded-For" Header Specifies if the proxy should send the X-Forwarded-For http header. Proxy Access Settings These settings configure the access method to your own http proxy and server. All traffic is routed through one single port, for both proxy and server. HTTPS Server Port: Server Access Restrictions You can restrict the access to this proxy/server using a two-stage security barrier: define an <em>access domain</em> with a list of granted client IP-numbers or with wildcards define an <em>user account</em> with an user:password - pair This is the account that restricts access to the proxy function. You probably don't want to share the proxy to the internet, so you should set the IP-Number Access Domain to a pattern that corresponds to you local intranet. The default setting should be right in most cases. If you want, you can also set a proxy account so that every proxy user must authenticate first, but this is rather unusual. IP-Number filter Accounts "'Referer' section from the standard IETF specification" "Link types section at W3C HTML specification" "Submit" Referrer Policy Settings When loading pages and navigating through links, a web browser sends some information about the origin of the request, Visited websites can process this information as they wish, so this can become a privacy concern, for example when coming from a page which contains searched terms in its URL. This page offers some configuration settings to instruct your browser how it should fill this referrer information. Beware that every browser behaves differently: some settings may be unsupported by your particular browser and therefore ignored. If you are really concerned about privacy, please check what is really sent by your browser by using its embedded developer tools network console, or with the network traffic analyzer of your choice. Global policy This referrer policy applies for every page on this peer. It is set by the "meta" HTML tag. Values are sorted by decreasing privacy level. no-referrer Highest privacy setting: referrer information should never be sent, even when navigating on this peer internal links. Be careful with this: some websites might reject requests with no referrer. same-origin Peer internal links: referrer information should be stripped from any private data and contain only this peer host name. External links: referrer information should never be sent. strict-origin Peer internal and external links: referrer information should be stripped from any private data and contain only this peer host name. Restriction: when a link downgrades from a TLS secured connection (https) on this peer to an unsecured target (http), no referrer information at all should be sent. origin strict-origin-when-cross-origin Peer internal links: referrer information should contain full URLs. External links: referrer information should be stripped from any private data and contain only this peer host name. Restriction: when an external link downgrades from a TLS secured connection (https) on this peer to an unsecured target (http), no referrer information at all should be sent. origin-when-cross-origin no-referrer-when-downgrade Referrer information should contain full URLs, except when a link downgrades from a TLS secured connection (https) on this peer to an unsecured target (http). empty value Default browser behavior: it should correspond to "no-referrer-when-downgrade". unsafe-url Unsafe setting: referrer information should always contain full URLs. Custom setting: probably manually edited, be sure this value is the desired one. Search results links Add the "noreferrer" link type to search results links When checked, this overrides the global referrer policy and adds the standard "noreferrer" thus instructing the browser that it should not send any referrer information at all when visiting them. It is a standard HTML5 attribute value, supported by many more browsers than the meta tag: if you want a higher level of privacy but use an old or incompatible browser, this can be a valuable option. <em id="submitInfo">Changes will take effect immediately.</em> "Submit" "Retry Uploading" Seed Upload Settings With these settings you can configure if you have an account on a public accessible server where you can host a seed-list file. General Settings: If you enable one of the available uploading methods, you will become a principal peer. Your peer will then upload the seed-bootstrap information periodically, but only if there have been changes to the seed-list. Upload Method Here you can specify which upload method should be used. Select 'none' to deactivate uploading. URL The URL that can be used to retrieve the uploaded seed file, like http://www.&lt;my-host&gt;.net/yacy/seed.txt' "Submit" Store into filesystem: You must configure this if you want to store the seed-list file onto the file system. File Location: Here you can specify the path within the filesystem where the seed-list file should be stored. current: "Submit" Uploading via FTP: This is the account for a FTP server where you can host a seed-list file. If you set this, you will become a principal peer. Your peer will then upload the seed-bootstrap information periodically, but only if there had been changes to the seed-list. Server The host where you have a FTP account, like 'ftp.&lt;my-host&gt;.net' Path The remote path on the FTP server, like 'yacy/seed.txt'. Missing sub-directories are NOT created automatically. Username Your log-in at the FTP server Password The password "Submit" Uploading via SCP: This is the account for a server where you are able to login via ssh. Server The host where you have an account, like 'my.host.net' Server&nbsp;Port The sshd port of the host, like '22' Path The remote path on the server, like '~/yacy/seed.txt'. Missing sub-directories are NOT created automatically. Username Your log-in at the server Password The password "Submit" Server Access Settings IP-Number filter: (requires restart) <strong>Here you can restrict access to the server.</strong> By default, the access is not limited, because this function is needed to spawn the p2p index-sharing function. If you block access to your server (setting anything else than '*'), then you will also be blocked from using other peers' indexes for search service. However, blocking access may be correct in enterprise environments where you only want to index your company's own web pages. Filter have to be entered as IP, IP range or using CIDR notation separated by comma (e.g. 192.168.1.1,2001:db8 ff00:42:8329,192.168.1.10-192.168.1.20,192.168.1.30-40,192.168.2.0/24) further details on format see Jetty staticIP (optional): <strong>The staticIP can help that your peer can be reached by other peers in case that your peer is behind a firewall or proxy.</strong> You can create a tunnel through the firewall/proxy (look out for 'tunneling through https proxy with connect command') and create an access point for incoming connections. This access address can be set here (either as IP number or domain name). If the address of outgoing connections is equal to the address of incoming connections, you don't need to set anything here, please leave it blank. If the value you enter here does not match with this IP, you will not be able to access the server pages anymore. publicPort (optional): <strong>The publicPort can help that your peer can be reached by other peers in case that your peer is behind a reverse proxy.</strong> If the port used to access YaCy is the same port the application is listening on, fileHost: Set this to avoid error-messages like 'proxy use not allowed / granted' on accessing your Peer by its hostname. Virtual host for httpdFileServlet access for example http://FILEHOST/ shall access the file servlet and return the defaultFile at rootPath either way, http://FILEHOST/ denotes the same as http://localhost:&lt;port&gt;/ for the preconfigured value 'localpeer', the URL is: http://localpeer/. Server Port Settings Server port: This is the main port for all http communication (default is 8090). A change requires a restart. Server ssl port: This is the port to connect via https (default is 8443). A change requires a restart. Shutdown port: This is the local port on the loopback address (127.0.0.1 or :1) to listen for a shutdown signal to stop the YaCy server (-1 disables the shutdown port, recommended default is 8005). A change requires a restart. Compression settings Compress responses with gzip When checked (default), HTTP responses can be compressed using gzip. The requesting user-agent (a web browser, another YaCy peer or any other tool) uses the header 'Accept-Encoding' to tell whether it accepts gzip compression or not. This adds some processing overhead, but can significantly reduce the amount of bytes transmitted over the network. <em id="submitInfo">Changes need a server restart.</em> "Submit" URL Proxy Settings With this settings you can activate or deactivate URL proxy. Service call: http://localhost:8090/proxy.html?url=parameter, where parameter is the url of an external web page. URL proxy: Enabled Globally enables or disables URL proxy via http://yourpeer:yourport/proxy.html?url=http://externalurl/ Show search results via URL proxy: Enables or disables URL proxy for all search results. If enabled, all search results will be tunneled through URL proxy. Alternatively you may add this javascript to your browser favorites/short-cuts, which will reload the current browser address via the YaCy proxy servlet. or right-click this link and add to favorites: Restrict URL proxy use: Define client filter. Default: 127.0.0.1,0:0:0:0:0:0:0:1. URL substitution: Define URL substitution rules which allow navigating in proxy environment. Possible values: all, domainlist. Default: domainlist. Advanced Settings If you want to restore all settings to the default values, but <strong>forgot your administration password</strong>, you must stop the proxy, delete the file 'DATA/SETTINGS/yacy.conf' in the YaCy application root folder and start YaCy again. Server Access Settings Referrer Policy Settings Crawler Settings Seed Upload Settings Message Forwarding (optional) Transparent Proxy Access Settings URL/Web Proxy Access Settings Remote Proxy (optional) Debug/Analysis Settings HTTP client Settings "Fork me on GitHub" "YaCy Websearch" "PerformanceGraph" "banner" "bad" "idea" "Update YaCy" "lock icon" "good" Log-in as administrator to see full status Welcome to YaCy! Your settings are _not_ protected! and set an administration password. You have not published your peer seed yet. This happens automatically, just wait. Your network configuration is in private mode. Your peer seed will not be published. Access is unrestricted from localhost (this includes administration features). The peer must go online to get a peer address. You cannot be reached from outside. A possible reason is that you are behind a firewall, NAT or Router. global index on your own search page. We encourage you to open your firewall for the port you configured (usually: 8090), or to set up a 'virtual server' in your router settings (often called DMZ). Please be fair, contribute your own index to the global index. it as soon as possible and restart YaCy. Crawling is paused! If the crawling was paused automatically, please check your disk space. You can download a more recent version of YaCy. Click here to install this update and restart YaCy: You are running a server in senior mode and you support the global internet index, You have a principal peer because you publish your seed-list to a public accessible server If you need professional support, please write to support@yacy.net System Status System Unknown Protection Default password is not changed [Configure] password-protected Address peer address not assigned Port Forwarding Host broken connected Proxy Transparent on off URL Remote: not used Yes No Auto-popup on start-up Tray-Icon Experimental Memory Usage RAM used: RAM max: DISK used: DISK free: Incoming Connections Queues Local Crawl (paused) Remote triggered Crawl Pre-Queueing Seed server Disabled. "Kaskelix" "Restart" "Shutdown" No action submitted Re-Start Shutdown Your system is not protected by a password You don't have the correct access right to perform this task. Please log in. See you soon! Application will terminate after working off all scheduled tasks. Please send us feed-back! We don't track YaCy users, YaCy does not send 'home-pings', we do not even know how many people use YaCy as their private search engine. Therefore we like to ask you: do you like YaCy? Will you use it again... if not, why? Is it possible that we change a bit to suit your needs? Please send us feed-back about your experience with an or a Professional Support Just a moment, please! Then YaCy will restart. If you can't reach YaCy's interface after 5 minutes restart failed. YaCy will be restarted after installation. <b>The file you are trying to install is not located in the release directory. You are in a development environment or the file you are trying to install is empty. "YaCy Supporter" "bookmark" "Add to bookmarks" "positive vote" "Give positive vote" "negative vote" "Give negative vote" Supporter Supporter are switched off for users without authorization "YaCy Surftips" "bookmark" "Add to bookmarks" "positive vote" "Give positive vote" "negative vote" "Give negative vote" "authentication required" Surftips Surftips are switched off for users without authorization YaCy Supporters a list of home pages of yacy users Show surftips to everyone Hide surftips for users without authorization "robots.txt Table" "API" The information that is presented on this page can also be retrieved as XML. Click the API icon to see the XML. robots.txt table "Tables" "Search" "Edit Selected Row" "Add a new Row" "Delete Selected Rows" "Delete Table" "Commit" Table Administration Table Selection Select Table: show max. all entries, reverse: search rows for PK Row Editor Primary Key "Single Threaddump" "Multiple Dump Statistic" YaCy Debugging: Thread Dump Threaddump Tools Add superpowers to the YaCy Chat. Tools may be disabled by setting maxCallsPerTurn to 0. Tool settings were saved. Basic Tools maxCallsPerTurn disable Visualization Tools Data Retrieval Tools Save Tools Configuration CyTag Trails "Publish" "negative vote" "positive vote" You can share your local addition to translations and distribute it to other peers. The remote peer can vote on your translation and add it to its own local translation. File: Originator English: existing Translation: Vote on this translation. If you vote positive the translation is added to your local translation list. "Save translation" Translation Editor Translate untranslated text of the user interface (current language). The modified translation file is stored in DATA/LOCALE directory. UI Translation Source File view it filter untranslated Source Text "login" "logout" "red bar" "green bar" "Change" User Page You are not logged in. Username: Password: (Identified by IP Username/Password Cookie old Password new Password new Password(repetition) You are currently logged in as admin. (after logout you will be prompted for your password again. simply click "cancel") Password was changed. Old Password is wrong. New Password and its repetition do not match. New Password is empty. "File system browser" "Root contents" Virtual File System User storage in the browser cache with file-system-like navigation. New Folder Upload File No files yet. Upload a file or create a folder. Preview Edit file Discard Save "API" "Show Metadata" "Browse Host" "Show Snippet" "Show" "action" See the page info about the url. View URL Content Get URL Viewer URL: Search in Document: URL Metadata Hash: In Metadata: no yes In Cache: First Seen: Word Count: Description: Size: MimeType: Collections: View as Original from Web Original from Cache Plain Text Parsed Text Parsed Sentences Parsed Tokens/Words Link List Schema Fields Citation Report Unable to find URL Entry in DB Invalid URL Unable to download resource content. Unable to parse resource content. Unsupported protocol. Snippet Headline Teaser Text Original Content from Web Parsed Content dc:title dc:creator dc:subject dc:description dc:publisher dc:format dc:identifier dc:source geo:lat &amp; geo:long nr type name link text rel Parsed Tokens CitationReport "refresh" Server Log reversed order regex terms Invalid regular expression filter. "vCard" "rdf:foaf" "Onlinestatus" Local Peer Profile: Remote Peer Profile: Wrong access of this page The requested peer is unknown or a potential peer. The profile can't be fetched. Name Nick Name Homepage eMail ICQ Jabber Yahoo! MSN Skype Comment vCard "API" "View" "Uniform Resource Locator" "Standard CSV field delimiter" "Create" "Submit" The information that is presented on this page can also be retrieved as XML Click the API icon to see the RDF Ontology definition for this vocabulary. Vocabulary Administration Vocabularies can be used to produce a search navigation. A vocabulary must be created before content is indexed. The vocabulary is used to annotate the indexed content with a reference to the object that is denoted by the term of the vocabulary. The object can be denoted by a url stub that, combined with the term, becomes the url for the object. Vocabulary Selection Vocabulary Name Vocabulary Production Please provide a CSV file path or <abbr title="Uniform Resource Locator">URL</abbr>. Empty Vocabulary Auto-Discover from file name from page title from page title (split) from page author Objectspace It is possible to produce a vocabulary out of the existing search index. This is done using a given 'objectspace' which you can enter as a URL Stub. This stub is used to find all matching URLs. If the remaining path from the matching URLs then denotes a single file, the file name is used as vocabulary term. This works best with wikis. Try to use a wiki url as objectspace path. Import from a csv file File Path or <abbr title="Uniform Resource Locator">URL</abbr> Start line (first has index 0) Column for Literals Synonyms no Synonyms Auto-Enrich with Synonyms from Stemming Library Read Column Column for Object Link (optional) (first has index 0, if unused set -1) Charset of Import File Column separator Comma ',' Semicolon ';' Vocabulary Editor File [automatically generated, not stored, cannot be edited] Size Namespace Predicate Prefix Is Facet? (If checked, this vocabulary is used for search facets. Not feasible for large vocabularies!) Match terms from Cleartext Linked data/Semantic web annotations Modify Delete Literal Object Link add clear table (remove all terms) delete vocabulary "API" "minus" "plus" "change" "WebStructurePicture" The data that is visualized here can also be retrieved in a XML file, which lists the reference relation between the domains. With a GET-property 'about' you get only reference relations about the host that you give in the argument field for 'about'. With a GET-property 'latest' you get a list of references that had been computed during the current run-time of YaCy, and with each next call only an update to the next list of references. Click the API icon to see the XML file. Web Structure Host List host depth nodes time size Background Color Text Line Pivot Dot Other Dot Dot-end "all" "admin" "Submit" "Preview" "Discard" "Show" "Compare" (only granted to admin) Index - Grant Write Access to Edit Author: Text: Preview No changes have been submitted so far! Index Subject Change Date Last Author Start Page Versions Compare version from with version from Error You can use Changes will be published as announcement on YaCyNews Wiki-Code This table contains a short description of the tags that can be used in the Wiki and several other servlets of YaCy. For a more detailed description visit the Code Description These tags create headlines. If a page has three or more headlines, a table of content will be created automatically. Headlines of level 1 will be ignored in the table of content. ''text''<br />'''text'''<br />'''''text''''' These tags create stressed texts. The first pair emphasizes the text (most browsers will display it in italics), the second one emphasizes it more strongly (i.e. bold) and the last tags create a combination of both. &lt;s&gt;text&lt;/s&gt; Text will be displayed struck through &lt;u&gt;text&lt;/u&gt; underlined text Lines will be indented. This tag is supposed to mark citations, but may as well be used for styling purposes. These tags create a numbered list. These tags create an unnumbered list. ;word 1:definition 1 ;word 2:definition 2 ;;word 3:definition 3 ;word 4:definition 4 These tags create a definition list. This tag creates a horizontal line. [[pagename]] [[pagename|description]] This tag creates links to other pages of the wiki. [url] [url description] This tag creates links to external websites. [[Image:url]] [[Image:url|alt text]] [[Image:url|align|alt text]] This tag displays an image, it can be aligned left, right or center. [[Youtube:id]] [[Vimeo:id]] This tag displays a Youtube or Vimeo video with the id specified and fixed width 425 pixels and height 350 pixels. i.e. use [[Youtube:QZsWG4-7Qfk]] to embed this video: https://www.youtube.com/watch?v=QZsWG4-7Qfk i.e. use [[Vimeo:32200946]] to embed this video: http://vimeo.com/32200946 ||row 1, col 1||row 1, col 2 ||row 2, col 1||row 2, col 2 These tags create a table, whereas the first marks the beginning of the table, the second starts a new line, the third and fourth each create a new cell in the line. The last displayed tag closes the table. &lt;pre&gt; text &lt;/pre&gt; A text between these tags will keep all the spaces and linebreaks in it. Great for ASCII-art and program code. text<br />&nbsp;text<br />text If a line starts with a space, it will be displayed in a non-proportional font. "YaCy-Logo" YaCy Firefox Search-Plugin Installation: Simply click on the link shown below to integrate the YaCy Firefox Search-Plugin into your browser. In Mozilla Firefox, you can the Search-Plugin via the search box on the toolbar.<br />In Mozilla (Seamonkey) you can access the Search-Plugin via the Sidebar or the Location Bar. Install the YaCy search plugin. Similar documents from different hosts: List of Cited filter cited sentences filter off List of other web pages with citations "Submit" File Upload This form can be used to upload a file and assign it to an url. Example usage is the direct attachment of a content management system to YaCy to push newly changed files directly to the YaCy indexer. File Count synchronous commit Files to process: File Number Data URL Collection Last-Modified Content-Type The following attributes are only used for media type content Media-Title Media-Keywords () Result for the recently submitted file(s). You can also submit the same form using the servlet push_p.json to get push confirmations in json format. count successall false true countsuccess countfail Item Success Message fail ok If you want to push again files, use this form to pre-define a number of upload forms: "Submit" File Share This form can be used to share a (index) file Files to process: Result for the recently submitted file(s). You can also submit the same form using the servlet share.json to get push confirmations in json format. successall false true countsuccess countfail Item URL Success Message fail ok If you want to push again files, use this form to pre-define a number of upload forms: "Table" "Edit Table" PK "API" This search result can also be retrieved as XML. Click the API icon to see an example call to the search rss API. Title Author Description Subject Publisher Contributor Date Type YaCy Identifier Identifier Language Collections Load Date Referrer Identifier Referrer URL Document size Number of Words Inbound Links (anchors) Outbound Links (anchors) Incoming Links (citation) Location "Compare" Websearch Comparison Left Search Engine Right Search Engine Search Result loading.... "Donate!" Please support our work on YaCy! Github Sponsors beneficial: 5 &euro; generous: 25 &euro; gracious: 50 &euro; "YaCy" "Search..." "Restart" "Shutdown" "Community" "Help" "Chat" "Search" Administration Toggle navigation Re-Start Shutdown Forum Help About This Page JavaScript information <i>external</i>&nbsp;&nbsp;&nbsp;YaCy Tutorials <i>external</i>&nbsp;&nbsp;&nbsp;Download YaCy <i>external</i>&nbsp;&nbsp;&nbsp;Community (Web Forums) <i>external</i>&nbsp;&nbsp;&nbsp;Git Repository Sponsor YaCy is free software, so we need the help of many to support the development.<br/><b>You</b> can help by joining a sponsoring plan: <i>external</i>&nbsp;&nbsp;&nbsp;<b>become a Github Sponsor</b> <i>external</i>&nbsp;&nbsp;&nbsp;<b>become a YaCy Patreon</b> Please help! We need financial help to move on with the development! Chat Search First Steps Use Case &amp; Account Grab a whole site Monitoring System Status Peer-to-Peer Network Index Browser Network Access Crawler Monitor Production Crawler AI Lab Automation YaCy Packs &amp; Import/Export Content Semantic Target Analysis Index Administration System Administration Filter &amp; Blacklists RAM/Disk Usage &amp; Updates Search Portal Integration Portal Configuration Portal Design Ranking and Heuristics "Log in to use extended search features" "Search Interfaces" "Help" "Administration" Toggle navigation Log in Search Interfaces<b class="caret"></b> <b class="caret"></b> Web Search File Search Compare Search Chat URL Viewer Example Calls to the Search API: <i>API</i>&nbsp;&nbsp;&nbsp;YaCy JSON <i>API</i>&nbsp;&nbsp;&nbsp;YaCy RSS/Opensearch <i>API</i>&nbsp;&nbsp;&nbsp;Solr RSS/Opensearch <i>API</i>&nbsp;&nbsp;&nbsp;Solr Default Core / JSON <i>API</i>&nbsp;&nbsp;&nbsp;Solr Default Core / XML <i>API</i>&nbsp;&nbsp;&nbsp;Solr Webgraph Core / XML About This Page YaCy Tutorials JavaScript information <i>external</i>&nbsp;&nbsp;&nbsp;Download YaCy <i>external</i>&nbsp;&nbsp;&nbsp;Community (Web Forums) <i>external</i>&nbsp;&nbsp;&nbsp;Git Repository <i>external</i>&nbsp;&nbsp;&nbsp;Bugtracker Administration &raquo; "Help" Toggle navigation Search Interfaces<b class="caret"></b> Web Search File Search Compare Search Chat URL Viewer Example Calls to the Search API: <i>API</i>&nbsp;&nbsp;&nbsp;YaCy JSON <i>API</i>&nbsp;&nbsp;&nbsp;YaCy RSS/Opensearch <i>API</i>&nbsp;&nbsp;&nbsp;Solr RSS/Opensearch <i>API</i>&nbsp;&nbsp;&nbsp;Solr Default Core / JSON <i>API</i>&nbsp;&nbsp;&nbsp;Solr Default Core / XML <i>API</i>&nbsp;&nbsp;&nbsp;Solr Webgraph Core / XML About This Page YaCy Tutorials JavaScript information <i>external</i>&nbsp;&nbsp;&nbsp;Download YaCy <i>external</i>&nbsp;&nbsp;&nbsp;Community (Web Forums) <i>external</i>&nbsp;&nbsp;&nbsp;Git Repository <i>external</i>&nbsp;&nbsp;&nbsp;Bugtracker Administration &raquo; AI Lab LLM Selection RAG Config Tools Config Log Reports AI Shield Chat Access Tracker Server Access Access Grid Incoming Requests Overview Incoming Requests Details All Connections Local Search Log Host Tracker Access Rate Limitations Remote Search Cookie Menu Incoming&nbsp;Cookies Outgoing&nbsp;Cookies Filter &amp; Blacklists Blacklist Administration Blacklist Cleaner Blacklist Test Import/Export Application Status System Status Processes Server Log Log Reports Thread Dump Concurrent Indexing Memory Usage Search Sequence Messages Overview Incoming&nbsp;News Processed&nbsp;News Outgoing&nbsp;News Published&nbsp;News Community Data Surftips Local Peer Wiki Bookmarks System Administration Advanced Settings Performance Settings of Busy Queues Viewer and administration for database tables Advanced Properties UI Translations Web Crawler Processing Monitor Crawler Loader Rejected URLs Queues Local Global Remote No-Load Crawler Steering Scheduler and Profile Editor robots.txt Monitor Crawl Results Overview (1) Receipts (2) Queries (3) DHT Transfer (4) Proxy Use (5) Local Crawling (6) Global Crawling (7) Pack Import Load Web Pages Site Crawling Parser Configuration Design Appearance Language Search Page Layout Index Administration URL Database Administration Index Deletion Index Sources &amp; Targets Solr Schema Editor Field Re-Indexing Reverse Word Index Content Analysis Advanced Crawler Crawler/Spider Crawl Start (Expert) Crawling of MediaWikis Crawling of phpBB3 Forums Network Harvesting Network Scanner Remote Crawling Scraping Proxy Autocrawl Content Export / Import YaCy Packs Pack Generator Pack Downloader Pack Manager Export Index Export Solr Dump Export/Import Import RSS OAI-PMH WARC ZIM JsonList Database Reader phpBB3 Database MediaWiki Dump RAM/Disk Usage &amp; Updates Performance Web Cache Download System Update Portal Configuration Generic Search Portal Search Box Anywhere User Profile Local robots.txt Publication Wiki Blog Ranking and Heuristics Solr Ranking Config RWI Ranking Config Heuristics Content Semantic Automated Annotation Auto-Annotation Vocabulary Editor Knowledge Loader Target Analysis Mass Crawl Check Regex Test Use Case &amp; Accounts Basic Configuration Accounts Network Configuration Web Visualization Index Browser Web Structure Image Collage forwarding forward to remote peer "Extend media search results (images, videos or applications specific) to pages including such medias (provides generally more results, but eventually less relevant)." "Strictly limit media search results (images, videos or applications specific) to indexed documents matching exactly the desired content domain." "Reference alpha-2 language codes list" Search Text Images Audio Video Applications more options... Results per page Resource the peer-to-peer network only the local index Prefer mask restrict on show all Constraints: only index pages Media search Extended Strict Query Operators restrictions inurl:&lt;phrase&gt; only urls with the &lt;phrase&gt; in the url inlink:&lt;phrase&gt; only urls with the &lt;phrase&gt; within outbound links of the document filetype:&lt;ext&gt; only urls with extension &lt;ext&gt; site:&lt;host&gt; only urls from host &lt;host&gt; author:&lt;author&gt; only pages with as-author-annotated &lt;author&gt; tld:&lt;tld&gt; only pages from top-level-domains &lt;tld&gt; on:&lt;date&gt; only pages with &lt;date&gt; in content from:&lt;date1&gt; to:&lt;date2&gt; only pages with a date between &lt;date1&gt; and &lt;date2&gt; in content keyword:&lt;phrase&gt; only pages with keyword anotation containing &lt;phrase&gt; /http only resources from http or https servers /ftp /smb /file spatial restrictions /location only documents having location metadata (geographical coordinates) /radius/&lt;latitude&gt;/&lt;longitude&gt;/&lt;distance&gt; only documents within a square zone embracing a circle of given radius (in decimal degrees) around the specified latitude and longitude (in decimal degrees) ranking modifier /date sort by date (latest first) /near multiple words shall appear near "" (doublequotes) /language/&lt;lang&gt; heuristics /heuristic add search results from external opensearch systems Search Navigation keyboard shortcuts next result page previous result page automatic result retrieval browser integration after searching, click-open on the default search engine in the upper right search field of your browser and select 'Add "YaCy Search.."' search as rss feed json search results for ajax developers: get the search rss feed and replace the '.rss' extension in the search result url with '.json' YaCy JavaScript license information YaCy JavaScript files license information Script License Source YaCy Bookmarks YaCy Portalsearch: "Download Java Plug-in" "Processing.org" domaingraph : Built with Processing This browser does not have a Java Plug-in. Get the latest Java Plug-in here. Built with Processing "login" Your Username/Password is wrong. Username Password YaCy: Error Message YaCy request: unspecified error not-yet-assigned error You don't have an active internet connection. Please go online. Could not load resource. The file is not available. Your Account is disabled for surfing. Did you mean: "add bookmark" YaCy stop proxy (Warning: secure target viewed over normal http) "retrieve" remote crawl fetch test Retrieve remote crawl url list Target Peer: select rss terminal "select all" "deselect all" "add" Add Items to Blacklist Unable to store the items into the blacklist file: File Error! Unable to fetch data from file. YaCy-Peer &quot; &quot; not found. URL &quot; &quot; not found or empty list. Wrong Invocation! Please invoke with sharedBlacklist.html?name=PeerName Parse Error! An error occured while parsing XML data. Please check if the XML is valid. Blacklist source: Blacklist target: Blacklist item "YaCy" "Download Java Plug-in" "PerformanceGraph" "WebStructurePicture" "The yacy Network" YaCy System Terminal Monitor &lt;Search Form&gt; &lt;Crawl Start&gt; &lt;Status Page&gt; &lt;Shutdown&gt; Event Terminal Image Terminal Domain Monitor This browser does not have a Java Plug-in. Get the latest Java Plug-in here. Resource Monitor Network Monitor "Attach search results by default" "Search" "Attach a file" "Send" "Clear chat" "Download chat" "Upload chat" "Show system prompt" YaCy Chat This Chat is private. YaCy does not keep any history — only your browser remembers the current conversation. Default Dialog Augmentation: no search, allow attachments use local search use global search User Attach Search Results Attach PNG/JPG or text (.txt/.md/.tex) Clear Chat Download Chat Upload Chat Show System "Search..." "Search" YaCy Interactive Search Click the API icon to see an example call to the search rss API. loading from local index... onkeyup="xmlhttpPost(); return false;" "Refresh sorting. Depending on their rank, some results fetched in background may then appear on this page." "YaCy server is fetching results from available data sources." "Show anyway links to images that could not be rendered" "Hide links to images that could not be rendered" "Play all" "Stop all" Click the RSS icon to see this search result as RSS message stream. Use the RSS search result format to add static searches to your RSS reader, if you use one. search No Results. No Results. (length of search words must be at least 1 character) You are not allowed to search the web with this peer. You have reached the maximum allowed number of accesses to this search page within ten minutes. Please try again later or log in as administrator or as a user with extended search right. You have reached the maximum allowed number of accesses to this search page within one minute. You have reached the maximum allowed number of accesses to this search page within three seconds. Did you mean: Location -- click on map to enlarge Failed to render <strong id="imageErrorsCount">0</strong> thumbnail(s). Show Hide Media URL Player "API" "search" The information that is presented on this page can also be retrieved as XML Click the API icon to see the XML. search "bookmark" "recommend" "delete" "blacklist host" "Show all" "Last known modification date" "Browse index" "Raw ranking score value" Tags: Metadata Parser Citations Pictures Cache View via proxy Not supported "Previous page" "Next page" &laquo; &raquo; "global" "local" "Use the default ranking profile (customizable), ordering results by score." "Use the 'Date' ranking profile, ordering results by default on each document last modification date." "text" "image" "audio" "video" "app" "false" "Extend media search results to pages including such medias (provides generally more results, but eventually less relevant)" "true" "Strictly limit media search results to indexed documents matching exactly the desired content domain." "earthsearchlogo" "Sorted by descending counts" "Sorted by ascending counts" "Sorted by descending labels" "Sorted by ascending labels" "click to expand facet" Peer-to-Peer Stealth Mode Privacy Stealth&nbsp;Mode Context Ranking Sort by Date Documents Images Audio Video Apps Extended Strict Location