Part II of The Art of Data War — Knowing Your Position, PRAISE Framework. Previously:
“We are not fit to lead an army on the march unless we are familiar with the face of the country: its mountains and forests, its pitfalls and precipices, its marshes and swamps.” ~ The Art of War, by Sun Tzu.
No firm decides to lose sight of its data. Very few firms treat data as a priority from the start; that is Prioritize’s territory, and it is why the problem arrives unannounced. Every team acquires its own vendors, builds its own datasets, and solves its own problem well while the firm grows. Then, at scale, data is suddenly a problem: nobody can say what the firm holds, what any of it means, or who is using it. A multi-strategy fund spends somewhere between tens and hundreds of millions of dollars a year on data, the second most expensive resource after people, and most could not produce the list. Getting the view back takes a data marketplace, under which every data asset has to earn its place, bought or built alike.
Visibility & Control, the Account dimension of PRAISE, assesses whether the organization has a clear and current view of its data landscape: what exists, what it represents, and how it’s being used. Organizations that lack visibility operate blind, making non-optimal decisions, duplicating effort, creating compliance risk, and wasting resources on data that delivers no value. Those that establish strong visibility and control can move faster, fuel innovation, govern effectively, and allocate resources where they actually matter.
Multi-Strategy Fund — Building Data Visibility at Scale
A large multi-strategy hedge fund faced a common problem as it scaled: portfolio management teams had no systematic way to discover what data already existed across the firm, and management could not tell which data assets were valuable. Teams operated in distinct pods across systematic strategies, equities long-short, and fixed income, each sourcing and processing data independently. A fund of that shape, spread across regions and asset classes, can easily hold a few hundred to a thousand datasets, and without a catalog tens of sourcing people do nothing but pass information, what we have and what to look for, while tens of engineers hand-build entitlements and manage feeds. The result was duplicated spending on licensing and effort, missed opportunities, and no institutional view of what the firm actually owned.
The firm built a centralized data catalog and teams changed how they worked with data. Each dataset was tagged with rich ontologies—sectors, asset classes, strategies—along with descriptive metadata covering collection methods, preparation workflows, and common use cases. This allowed portfolio managers to browse intelligently and compare alternative datasets before committing resources. It also connected sourcing strategy teams, licensing, and vendor management into a unified workflow, eliminating the friction that previously slowed data acquisition.
Once a dataset could be found, the next question was whether it was worth using, so the data management platform integrated dataset profiling metadata directly into the catalog. Teams could see quality metrics, completeness, historical depth, consistency, and timeliness before ever requesting access. For systematic teams especially, they could now eliminate unsuitable datasets in hours rather than spending weeks or months on manual exploration. Data lineage tracking revealed how datasets were generated and which qualification rules they passed, adding another layer of confidence.
With integrated access controls and usage analytics, the catalog became a bidirectional marketplace: data went out to serve users, and evidence of what it was worth came back to the owners. Data sourcing teams could now measure the actual value of datasets based on who used them and how often, identify synergies across pods, and negotiate renewals from strength. The people who had spent their days passing information were freed to add value: finding new sources of alpha, understanding what portfolio managers needed, and building relationships with vendors. The organization stopped treating data as infrastructure and started managing it as a strategic asset, the shift Prioritize describes. Every data asset, bought or built, now had to earn its place, and for the first time the firm could say which ones did.
The Signals to Look For
Read the case back and three things carried it. The firm could find what it had, because every dataset was described the way the business thinks. It could judge a dataset before committing resources to it, because the profiling, lineage and other important meta/operational data sat on the listing. And it could see what each dataset was used for and what it was worth, because access and usage ran through the same place. Those are the three questions to put to your own organization, in that order. How to build one is the subject of Platform as a Foundation, later in the book.
1. Do you know what data assets you have and where?
What to look for: Does your organization maintain a catalog that captures what data exists, where it lives, who owns it, what it represents, and how to access it—kept up to date as part of standard data onboarding and engineering processes? Or does discovering available data require tribal knowledge, emails to multiple teams, and luck?
The most fundamental requirement for data visibility is knowing what you have. This sounds obvious, but in practice, most organizations have no systematic answer. Data lives across cloud platforms, on-premise systems, third-party vendors, departmental databases, and individual file shares. Teams acquire datasets, create new ones, or enrich existing data independently, and no one maintains a central view.
For organizations that depend on a high number of diverse datasets for intelligence, investment firms, research institutions, sales organizations, the absence of a catalog creates severe competitive disadvantage. Some analysts are completely blocked from datasets they don’t know exist. Others have technical access to tables or files but lack context about what the data represents, its quality, or appropriate use cases, leading to misinterpretation or underutilization. Teams spend hours on calls trying to piece together what’s available. Datasets the firm built get built again: a derived dataset or a signal one desk produced is rebuilt by another because no one knows it exists, duplicating the effort and often producing two versions of different quality, and somewhere a portfolio manager is investing on the weaker one. Worst of all, teams acquire licenses for datasets the firm already owns or has a comparable alternative, wasting capital on duplicates. The impact extends to recruiting: new portfolio managers find the move less attractive when they can’t efficiently discover and access data to support their alpha research.
A data marketplace solves these problems by creating a searchable, browsable catalog of datasets described the way the business thinks. An entry describes what the data represents, the business domains it belongs to, its common use cases, and its update frequency and scope, and is tagged with ontologies—asset classes, customer segments, product lines, geographies—so users can explore intelligently. AI search now lets users find datasets by asking in plain language, matching on metadata, semantic relationships and usage patterns.
This is a different product from the catalog most firms buy. Tooling providers like Alation and Collibra offer platforms to build and manage catalogs, and their unit is the table, the column, and the schema: a technical catalog, organized around how a database is structured. Business glossaries and metadata sit on top, but the thing being catalogued is still the table. They struggle to represent data at the level business users actually think about it: coherent datasets that answer specific questions, not fragmented across dozens of tables.
Catalog freshness matters as much as coverage. A catalog is only valuable if it’s current—stale catalogs stop being trusted and stop being used. This means catalog updates must be integrated into standard data lifecycle processes, not treated as a separate responsibility. In a firm that has this, engineering teams update the catalog as part of any deployment, and sourcing teams maintain the entry as part of a license renewal. Where the catalog is a separate duty, it is already out of date.
Ownership is equally critical. Every data asset needs a clear owner. Accountability is one reason; the other is that ownership is what keeps the catalog reliable and the marketplace working. Owners are responsible for keeping metadata current, which makes the catalog trustworthy. The owner might be a business function releasing operational data—sales teams publishing end-of-day numbers, pharmaceutical research teams sharing clinical trial results. It might be a sourcing team managing vendor relationships and ensuring product details are reflected in the catalog. In some cases, external vendors serve as owners for third-party data they provide directly.
External catalog providers exist because organizations struggle to discover relevant data sources beyond their walls. Firms like Eagle Alpha and Neudata help investment organizations explore alternative datasets across the market—tracking new vendors, evaluating emerging data products, and understanding what competitors might be using. Integrating with them is valuable. But external catalogs can never fully replace internal ones. They lack knowledge of your proprietary datasets, your internal enrichment processes, your specific use cases, and the institutional context that makes data valuable in your environment.
Knowing what you have is the first question, and the marketplace answers it. The harder two follow: whether a dataset is worth using, and whether it is being used.
2. Do you know what your datasets actually represent and when they’re valuable?
What to look for: Beyond knowing datasets exist, can users assess whether a dataset is fit for their specific purpose before investing time exploring it? Can they see quality metrics, coverage, timeliness, lineage, where the dataset is available, how well it is supported, and what it costs and permits? Or does evaluation require manual exploration, conversations with data owners, and trial-and-error testing?
Discovering that a dataset exists is only the beginning. The catalog tells you where something is—assessment tells you whether it’s worth using. This is the difference between browsing Amazon product listings and actually reading the detailed descriptions, customer ratings, reviews, and comparing alternatives before you buy, where the listing is the dataset. Without integrated assessment capabilities, teams waste enormous time on datasets that turn out to be incomplete, stale, poor quality, or simply the wrong fit for their use case. What makes that assessment possible is the operational and the meta data the platform keeps on the listing.
Impactful catalogs are living systems, continuously enriched by the data management infrastructure that produces and maintains the datasets. Lineage is critical: analysts spend hours, sometimes days, tracing how they got specific values when lineage is unclear. The catalog should automatically capture and allow users to trace data sources, code versions, quality rules executed, and transformations applied. Coverage, completeness, consistency, uniqueness and validity metrics can be generated through automated profiling metadata or data quality rules wired into each dataset refresh. Timeliness becomes transparent by surfacing both processing times and SLA breaches—how often does this dataset arrive late? Channeling across downstream systems shows where the dataset actually lives, and which analytics platforms can reach it, so users know immediately whether they can access it in their preferred environment. Operational metadata also reveals the level and quality of support the dataset receives from its owners, signaling how actively it’s maintained.
The metadata the platform produces is only half of what makes a listing trustworthy. The other half comes from the users themselves: feedback, ratings, use case descriptions from people who’ve actually worked with the data. That is what makes the catalog a marketplace, and the challenge is keeping it current. Smart organizations create incentives, and the incentive is an exchange: a new dataset gets a trial period; production access is granted against a documented use case and an assessment of whether the dataset improved the results; and a failed trial is closed with the feedback that earns the next one. Users pay for access in the currency the marketplace needs most.
The marketplace surfaces pricing and compliance rules the same way. Owners can tag licensing costs, or price them dynamically, whether to allocate ownership costs across teams or to price their own services. Compliance teams flag restrictions: headcount limits, prohibited use cases, AI usage constraints, geographic limitations. This information is essential for access control gateways, but it’s equally valuable to socialize upfront so users don’t waste time exploring datasets they’re not permitted to use.
The value of this marketplace model is speed. Analysts explore more datasets because they can fail fast, assessing that something isn’t appropriate and moving on, or proceed with confidence knowing quality, coverage, and compliance align with their needs. Without this, teams spend weeks gaining access, normalizing data, and testing—only to discover at the end that the dataset won’t work. That delay kills innovation. When assessment is integrated, teams iterate faster, test more hypotheses, and spend their time on datasets that will hold.
3. Do you understand the value and ROI of your data assets?
What to look for: Can you see which data assets are actively delivering value versus sitting idle, and what each of them costs? Do you know who’s using them and for what purpose? Does this visibility surface whether usage is appropriate and compliant? Or is usage data scattered across systems and invisible?
Datasets are distributed through various channels—databases, file systems, APIs—each with their own access control mechanisms. But access shouldn’t be managed purely as a technical concern by data engineering teams. Owners or their delegates, such as sourcing teams managing vendor data, should have the levers to control permissions directly through the marketplace catalog. The catalog maintains logs of who requested access and when it was granted. Before approving access, owners verify rules and restrictions, and kick off pricing approvals if necessary. This shifts access control from infrastructure administration to business-governed processes, ensuring permissions align with data policy rather than just technical capability.
Granting permission isn’t the same as understanding usage. To assess value, you need visibility into actual consumption: who’s querying the data, how often, and for what purpose. Usage telemetry captures which schemas or subsets are accessed, what filters are applied, and whether the data supports exploratory analysis or production systems, and the distribution channels integrate back to the data management platform to create a unified usage model.
Consumption is one half of the reading. The other half is cost, and not all catalog metadata is meant for end users—some exists purely for governance and control. Total cost of ownership (TCO) is a prime example. The catalog tracks not just licensing fees, but the full economic footprint of each data asset: storage costs, compute resources consumed during processing and distribution, engineering time spent maintaining pipelines, and operational overhead for support and troubleshooting. This gives data owners and finance teams visibility into a data asset’s full cost burden, and it is the number consumption has to be read against.
Put the two together and the catalog answers the question its heading asks. For any data asset, bought or built, the firm can set what it is used for and by whom against what it costs to keep, and say whether it earns its place: keep it, expand it, or retire it. A team that knows what a data asset is worth negotiates its renewal from strength. It also sees whether usage aligns with intended purposes and licensing terms. How the entitlement and usage data the marketplace collects feed ROI assessments, chargebacks and pricing is the Value Framework’s subject, later in the book; the reading itself is one a firm can produce today.
The same visibility is what an audit tests. Data vendors, government agencies and regulatory bodies audit how data is used, and responding can consume enormous time and resources when visibility is poor. Slow or incomplete responses not only waste effort—they signal to auditors that your house isn’t in order, which can trigger deeper, more invasive audits. When you have integrated visibility—catalog entries connected to lineage details, derived datasets, compliance rules, access controls, and usage patterns—audit responses become fast and confident. You can demonstrate exactly who accessed what data, for which purposes, under what permissions, and whether usage complied with licensing terms. Compliance shifts from a burden to an operational advantage, and auditors recognize an organization that takes data governance seriously.
Now read your own organization. Can an analyst find a dataset without emailing anyone? Can they judge whether it fits before requesting access? Can you say, for any data asset you bought or built, whether it earns what it costs? Three yeses mean the firm has strong marketplace dynamics, whatever it calls it. Three noes mean it has an inventory of tables and a licensing bill it cannot explain. The first question is the easiest to fix and the third is the one that decides whether the catalog is operating or merely built.



