Your data

Articles and research

Records and settings for blogs, news, knowledge bases, and research libraries: titles and authors first, long text within the limits, and sorting by date or popularity.

Readers search articles by title, author, and topic, then narrow by date. This recipe sets up blogs, news archives, knowledge bases, and research libraries for that. The papers playground searches 9,000 research papers set up this way.

Records

One record per article. A paper from the playground:

{
  "objectID": "1706.03762",
  "title": "Attention Is All You Need",
  "authors": [
    "Ashish Vaswani",
    "Noam Shazeer",
    "Niki Parmar",
    "Jakob Uszkoreit",
    "Llion Jones",
    "Aidan N. Gomez",
    "Lukasz Kaiser",
    "Illia Polosukhin"
  ],
  "authorCount": 8,
  "abstract": "The dominant sequence transduction models are based on complex recurrent…",
  "categories": ["cs.CL", "cs.LG"],
  "year": 2017,
  "published": 1497290254,
  "citations": 26678
}

A blog post or news story has the same shape: title, authors, a summary or the opening of the body, tags, published, a popularity number such as views, and the url and image results show.

  • Long text. A searchable field holds up to 16 KB of text. For longer articles, search the title, summary, and opening, or make one record per section as for docs, which shows an article once for each section that matches.
  • Authors are a list. As the second searchable field, they hold 1 KB: for long author lists, keep the first authors and store how many there are, like authorCount.
  • Dates twice: published in Unix seconds to sort by and filter exact ranges, and year for a year filter or facet.
  • Topics are a list, such as categories or tags, for facets.
  • Popularity is a number, such as views, shares, or citations.

Settings

{
  searchableAttributes: [
    { field: "title", weight: 8 },
    { field: "authors", weight: 3 },
    { field: "abstract", weight: 1 },
  ],
  facetFields: ["categories"],
  filterFields: ["year"],
  sorts: [
    { id: "relevance", label: "Relevance", field: "relevance", direction: "desc", thenBy: [] },
    { id: "cited", label: "Most cited", field: "citations", direction: "desc", thenBy: [] },
    { id: "newest", label: "Newest", field: "published", direction: "desc", thenBy: [] },
  ],
  defaultSort: "relevance",
  customRanking: [{ field: "citations", direction: "desc" }],
  synonyms: [["llm", "large language model"], ["rag", "retrieval-augmented generation"]],
  didYouMean: true,
  languages: ["en"],
}
  • Acronyms are synonyms, so "LLM" finds "large language model" and the other way round.
  • Popularity only settles near-ties. Custom ranking orders results within 10% of the best match's score, so a famous article never outranks a much better match. For the most read or most cited first, offer a sort, as above.
  • With no query, results follow custom ranking, so an empty search box shows the most cited papers first.
const results = await client.search({
  q: "retrieval augmented generation",
  filters: { year: { gte: 2020 } },
  facets: ["categories"],
  sort: "cited",
  highlight: true,
})

Show the matching part of a long field, such as the abstract, with Snippet (see Highlights).