问题:Elasticsearch - Java RestHighLevelClient - 如何使用滚动api获取所有文档

在我的 Elasticsearch 索引中,我保存了大约 30000 个实体。我想使用 RestHighLevelClient 获取它们的所有 id。我读过最好的方法是使用滚动 api。但是,当我这样做时,我只收到大约 10 个实体而不是 30k。如何解决这个问题

final class ElasticRepo {
    private final RestHighLevelClient restHighLevelClient;

List<ListingsData> getAllListingsDataIds() {
        val request = new SearchRequest(ELASTICSEARCH_LISTINGS_INDEX);
        request.types(ELASTICSEARCH_TYPE);
        val searchSourceBuilder = new SearchSourceBuilder()
                .query(matchAllQuery())
                .fetchSource(new String[]{"listing_id"}, new String[]{"backoffice_data", "search_and_match_data"});
        request.source(searchSourceBuilder);
        request.scroll(TimeValue.timeValueMinutes(3));
        return executeQuery(request);
    }

 private List<ListingsData> executeQuery(final SearchRequest searchQuery) {
        try {
            val hits = restHighLevelClient.search(searchQuery, RequestOptions.DEFAULT).getHits().getHits();
            return Arrays.stream(hits).map(SearchHit::getSourceAsString).map(ElasticRepo::toListingsData).collect(Collectors.toList());
        } catch (IOException e) {
            e.printStackTrace();
            throw new RuntimeException("");
        }
    }

}

当我这样做时,executeQuery 只返回大约 11 个实体。如何解决,如何获取索引中的所有文档?

解答

尝试遵循此示例,我正在使用此代码并且它有效:

        String query = "your query here";

        QueryBuilder matchQueryBuilder = QueryBuilders.boolQuery().must(new QueryStringQueryBuilder(query));

        SearchSourceBuilder searchSourceBuilder = new SearchSourceBuilder();

        searchSourceBuilder.query(matchQueryBuilder);

        searchSourceBuilder.size(5000); //max is 10000

        searchRequest.indices("your index here");

        searchRequest.source(searchSourceBuilder);

        final Scroll scroll = new Scroll(TimeValue.timeValueMinutes(10L));

        searchRequest.scroll(scroll);

        SearchResponse searchResponse = client.search(searchRequest);
            String scrollId = searchResponse.getScrollId();

        SearchHit[] allHits = new SearchHit[0];

        SearchHit[] searchHits = searchResponse.getHits().getHits();

        while (searchHits != null && searchHits.length > 0)
        {

            allHits = Helper.concatenate(allHits, searchResponse.getHits().getHits()); //create a function which concatenate two arrays

            SearchScrollRequest scrollRequest = new SearchScrollRequest(scrollId);

            scrollRequest.scroll(scroll);

            searchResponse = client.searchScroll(scrollRequest);

            scrollId = searchResponse.getScrollId();

            searchHits = searchResponse.getHits().getHits();

        }

        ClearScrollRequest clearScrollRequest = new ClearScrollRequest();
        clearScrollRequest.addScrollId(scrollId);
        ClearScrollResponse clearScrollResponse = client.clearScroll(clearScrollRequest);
Logo

欢迎大家访问Elastic 中国社区。由Elastic 资深布道师,Elastic 认证工程师,认证分析师,认证可观测性工程师运营管理。

更多推荐