<feed xmlns='http://www.w3.org/2005/Atom'>
<title>yacy/defaults/solr.webgraph.schema, branch master</title>
<subtitle>YaCy search server with some extra patches
</subtitle>
<id>https://git.fennell.dev/yacy/atom?h=master</id>
<link rel='self' href='https://git.fennell.dev/yacy/atom?h=master'/>
<link rel='alternate' type='text/html' href='https://git.fennell.dev/yacy/'/>
<updated>2014-08-01T11:20:25Z</updated>
<entry>
<title>target linktexts must be string to enable search facets on these fields</title>
<updated>2014-08-01T11:20:25Z</updated>
<author>
<name>orbiter</name>
<email>mc@yacy.net</email>
</author>
<published>2014-08-01T11:20:25Z</published>
<link rel='alternate' type='text/html' href='https://git.fennell.dev/yacy/commit/?id=2371d6b8dbdf5addf4d7d9124d88f2d976f8dab8'/>
<id>urn:sha1:2371d6b8dbdf5addf4d7d9124d88f2d976f8dab8</id>
<content type='text'>
</content>
</entry>
<entry>
<title>removed clickdepth_i field and related postprocessing. This information</title>
<updated>2014-04-16T20:16:20Z</updated>
<author>
<name>Michael Peter Christen</name>
<email>mc@yacy.net</email>
</author>
<published>2014-04-16T20:16:20Z</published>
<link rel='alternate' type='text/html' href='https://git.fennell.dev/yacy/commit/?id=9a5ab4e2c141ac9ef9f45c6524f2352ab7c3789c'/>
<id>urn:sha1:9a5ab4e2c141ac9ef9f45c6524f2352ab7c3789c</id>
<content type='text'>
is now available in the crawldepth_i field which is identical to
clickdepth_i because of a specific crawler strategy.</content>
</entry>
<entry>
<title>- reduce computation in case that specific postprocessing fields are not</title>
<updated>2013-12-04T16:48:12Z</updated>
<author>
<name>Michael Peter Christen</name>
<email>mc@yacy.net</email>
</author>
<published>2013-12-04T16:48:12Z</published>
<link rel='alternate' type='text/html' href='https://git.fennell.dev/yacy/commit/?id=e3c2f09de9fe8eb5db9d51006c67d474267ad045'/>
<id>urn:sha1:e3c2f09de9fe8eb5db9d51006c67d474267ad045</id>
<content type='text'>
selected
- de-select citation rank computation</content>
</entry>
<entry>
<title>added two more fields source_cr_host_norm_i,target_cr_host_norm_i in</title>
<updated>2013-09-27T14:57:05Z</updated>
<author>
<name>Michael Peter Christen</name>
<email>mc@yacy.net</email>
</author>
<published>2013-09-27T14:57:05Z</published>
<link rel='alternate' type='text/html' href='https://git.fennell.dev/yacy/commit/?id=b28d43decc29780eb9f60089cb47ecb3bbc856f7'/>
<id>urn:sha1:b28d43decc29780eb9f60089cb47ecb3bbc856f7</id>
<content type='text'>
webgraph and an addition to postprocessing to copy all cr ranking
attributes to the link edges associated to the postprocessing documents</content>
</entry>
<entry>
<title>added the new field harvestkey_s to the collection index and the</title>
<updated>2013-09-25T12:38:24Z</updated>
<author>
<name>Michael Peter Christen</name>
<email>mc@yacy.net</email>
</author>
<published>2013-09-25T12:38:24Z</published>
<link rel='alternate' type='text/html' href='https://git.fennell.dev/yacy/commit/?id=4f83d5f18c6fdcb7d007ba0495eaea957ce18362'/>
<id>urn:sha1:4f83d5f18c6fdcb7d007ba0495eaea957ce18362</id>
<content type='text'>
webgraph index which is temporary filled with the crawl profile key.
This is used to select a set of documents for post-processing as soon as
a crawl is finished. Now the postprocessing for a specific crawl is
started when that specific crawl is finished and not at the end of all
post-processing steps.</content>
</entry>
<entry>
<title>- replaced the properties object in AnchorURL with distinct variables</title>
<updated>2013-09-15T21:27:04Z</updated>
<author>
<name>Michael Peter Christen</name>
<email>mc@yacy.net</email>
</author>
<published>2013-09-15T21:27:04Z</published>
<link rel='alternate' type='text/html' href='https://git.fennell.dev/yacy/commit/?id=61c5e4068743dc0c8f8c9fa6e5e38be31b93e90a'/>
<id>urn:sha1:61c5e4068743dc0c8f8c9fa6e5e38be31b93e90a</id>
<content type='text'>
for anchor attributes.
- this caused that large portions of the parser code had to be adopted
as well
- added a counter target_order_i for anchor links in webgraph
computation
</content>
</entry>
<entry>
<title>added url_file_name_s in default collection schema for the file name</title>
<updated>2013-06-25T14:27:20Z</updated>
<author>
<name>Michael Peter Christen</name>
<email>mc@yacy.net</email>
</author>
<published>2013-06-25T14:27:20Z</published>
<link rel='alternate' type='text/html' href='https://git.fennell.dev/yacy/commit/?id=16d1d744faa4c1bf82403c852c273e70b24179a8'/>
<id>urn:sha1:16d1d744faa4c1bf82403c852c273e70b24179a8</id>
<content type='text'>
without the file extension. This part of the file path is removed from
the multi-field url_paths_sxt, which has now not the file name as last
part of the path list.

The same applies to the new fields source_file_name_s and
target_file_name_s in the webgraph schema.</content>
</entry>
<entry>
<title>increased number of links limitation from 1000 to 10000 for rss feeds</title>
<updated>2013-03-17T21:13:56Z</updated>
<author>
<name>orbiter</name>
<email>mc@yacy.net</email>
</author>
<published>2013-03-17T21:13:56Z</published>
<link rel='alternate' type='text/html' href='https://git.fennell.dev/yacy/commit/?id=17ae51e741abb7d60593a2e7e378895324e67457'/>
<id>urn:sha1:17ae51e741abb7d60593a2e7e378895324e67457</id>
<content type='text'>
and html documents</content>
</entry>
<entry>
<title>added clickdepth field writing for webgraph core (unfinished)</title>
<updated>2013-03-14T00:35:38Z</updated>
<author>
<name>orbiter</name>
<email>mc@yacy.net</email>
</author>
<published>2013-03-14T00:35:38Z</published>
<link rel='alternate' type='text/html' href='https://git.fennell.dev/yacy/commit/?id=6b13dd0d3dedd467765098e6382806cf6ece13e0'/>
<id>urn:sha1:6b13dd0d3dedd467765098e6382806cf6ece13e0</id>
<content type='text'>
</content>
</entry>
<entry>
<title>changes in ranking computation</title>
<updated>2013-03-13T13:47:00Z</updated>
<author>
<name>Michael Peter Christen</name>
<email>mc@yacy.net</email>
</author>
<published>2013-03-13T13:47:00Z</published>
<link rel='alternate' type='text/html' href='https://git.fennell.dev/yacy/commit/?id=addba047e297a8be1be17d368146abd22a0edac0'/>
<id>urn:sha1:addba047e297a8be1be17d368146abd22a0edac0</id>
<content type='text'>
- an existing ranking servlet for solr was extended. It is now possible
to set boost values for fields, boost functions and boost queries.
- The ranking can have different instances, but currently only the first
one is used
- added an abstraction layer for fields which can be used for search and
those fields can be edited in the solr ranking configruation
- the ranking value from solr within the field score is used to combine
remote search requests, which all are created using the same locally
defined boost values
- reduced the number of fields which are used for search (makes it
faster)
- replaced some text fields by string fields (makes indexing faster)
- removed classes which had no use
- made a large number of experiments for a better ranking and created a
temporary setting which prefers hits inside titles
- adjusted also the RWI-based ranking computation to 'prefer title'
- made special cases like for portal search where no post-processing and
post-ranking is wanted: this keeps the original ranking order as done by
Solr
- fixed many bugs with old settings for ranking</content>
</entry>
</feed>
