KonstantinSeurer/mesa

Commit Graph

Author	SHA1	Message	Date
Dave Airlie	f8d5b377c8	radv: set cb base tile swizzles for MRT speedups (v4) This patch uses addrlib to workout the tile swizzles according to the surface index. It seems to produce the same values as amdgpu-pro for the deferred test. v2: don't apply swizzle to CMASK. the eg docs don't mention it, and we clearly don't align cmask for that. v3: disable surf index for dedicated images, as these will most likely be shared, and I don't think the metadata has space for this info in it yet. v4: update for shareable images, rename combined_swizzle to tile_swizzle This gets the deferred demo from 730->950fps on my rx480. (dcc cmask elim predication patches get it further) Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Signed-off-by: Dave Airlie <airlied@redhat.com>	2017-07-17 01:43:41 +01:00
Dave Airlie	b86f86f55c	radv: allow clear merging for depth/stencil with no care stencil Some of the Sascha Willems demos pick a D32/S8 format for the depth buffer, then do a LOAD_OP_CLEAR/LOAD_OP_DONT_CARE on it, which means we don't get to merge the undefined->depth and clear htile transitions. This add the stencil aspect to the pending clears if there is a depth clear pending and the stencil aspect is don't care. Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Signed-off-by: Dave Airlie <airlied@redhat.com>	2017-07-17 01:16:59 +01:00
Bas Nieuwenhuizen	373f707fbb	radv: Remove NV dedicated alloc extension. To not confuse apps in thinking it might be faster. Signed-off-by: Bas Nieuwenhuizen <basni@google.com> Reviewed-by: Andres Rodriguez <andresx7@gmail.com>	2017-07-15 20:10:43 +02:00
Bas Nieuwenhuizen	515da29360	radv: Use the KHR dedicated alloc for the WSI. NV isn't valid for external images anymore. Signed-off-by: Bas Nieuwenhuizen <basni@google.com> Fixes: `6ddc64b93e` "radv: Add support for VK_KHR_dedicated_allocation." Reviewed-by: Andres Rodriguez <andresx7@gmail.com>	2017-07-15 20:10:25 +02:00
Jason Ekstrand	b70829708a	radv: Implement VK_KHR_external_memory This effectively reverts commit 43a171878bb4b5aedb36a. Technically, VK_KHR_get_memory_requirements2 and VK_KHR_dedicated_allocation are required for the KHR version but this at least restores the removed functionality. This patch builds but has received zero testing. Acked-by: Dave Airlie <airlied@redhat.com>	2017-07-15 08:59:38 -07:00
Bas Nieuwenhuizen	6ddc64b93e	radv: Add support for VK_KHR_dedicated_allocation. Signed-off-by: Bas Nieuwenhuizen <basni@google.com> Reviewed-by: Jason Ekstrand <jason@jlekstrand.net> Acked-by: Dave Airlie <airlied@redhat.com>	2017-07-15 08:59:38 -07:00
Bas Nieuwenhuizen	97931f0297	radv: Add support for VK_KHR_get_memory_requirements2. Fished the SparseImage call out of the headers as the spec missed the definition. Signed-off-by: Bas Nieuwenhuizen <basni@google.com> Reviewed-by: Jason Ekstrand <jason@jlekstrand.net> Acked-by: Dave Airlie <airlied@redhat.com>	2017-07-15 08:59:38 -07:00
Jason Ekstrand	0ee8d81718	anv: Implement VK_KHR_external_memory_* Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com> Reviewed-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com>	2017-07-15 08:59:38 -07:00
Jason Ekstrand	c02da9cad6	anv: Implement VK_KHR_dedicated_allocation We always recommend sub-allocation and don't do anything special for dedicated allocations. Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com> Reviewed-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com>	2017-07-15 08:59:38 -07:00
Jason Ekstrand	8c82aa5f43	anv: Implement VK_KHR_get_memory_requirements2 Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com> Reviewed-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com>	2017-07-15 08:59:38 -07:00
Jason Ekstrand	5b57bdc1cf	anv: Advertise version 1.0.54 Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com> Reviewed-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com>	2017-07-15 08:59:38 -07:00
Jason Ekstrand	227debdc92	vulkan: Update to the new 1.0.54 spec XML and headers There is one small ANV change here because we used the VK_ERROR_INVALID_EXTERNAL_HANDLE_KHX enum in the BO cache and that had to be updated to have the _KHR suffix. Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com> Reviewed-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com>	2017-07-15 08:59:38 -07:00
Jason Ekstrand	3b95e03b2c	radv: Drop support for VK_KHX_external_semaphore_* These have been formally deprecated by Khronos never to be shipped again. The KHR versions should be implemented/used instead. Acked-by: Dave Airlie <airlied@redhat.com>	2017-07-15 08:58:55 -07:00
Jason Ekstrand	dc179aa123	anv: Drop support for VK_KHX_external_semaphore_* These have been formally deprecated by Khronos never to be shipped again. The KHR versions should be implemented/used instead. Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com> Reviewed-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com>	2017-07-15 08:58:51 -07:00
Jason Ekstrand	4ac94d0dee	anv: Drop support for VK_KHX_external_memory_* These have been formally deprecated by Khronos never to be shipped again. The KHR versions should be implemented/used instead. Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com> Reviewed-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com>	2017-07-14 22:12:39 -07:00
Matt Turner	5ffe0c9e1b	i965: Compile with -msse2 (instead of -msse2) Ian noted that were were two Pentium 4 Extreme Edition LGA 775 CPUs, and they only have SSE2.	2017-07-14 22:01:11 -07:00
Matt Turner	6b05c080f2	i965: Compile with -msse3 All CPUs that can be paired with a GPU supported by i965_dri.so supports SSE3. This allows us to ensure that some vectorized version of the tiled memcpy path is enabled on 32-bit systems. This also ensures that __builtin_ia32_clflush is always usable. Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=101774 Tested-by: Tobias Klausmann <tobias.johannes.klausmann@mni.thm.de> Acked-by: Kenneth Graunke <kenneth@whitecape.org>	2017-07-14 16:54:43 -07:00
Kenneth Graunke	c7af6d2690	egl: Fix predecence problem when setting __DRI_CTX_FLAG_NO_ERROR This accidentally set __DRI_CTX_FLAG_NO_ERROR whenever any flags were present. Just needs extra parenthesis. Fixes: `4909519a66` (egl: Add EGL_KHR_create_context_no_error support) Reviewed-by: Grigori Goronzy <greg@chown.ath.cx> Tested-by: Mark Janes <mark.a.janes@intel.com>	2017-07-14 15:25:44 -07:00
Petr Sebor	b317cc1b3c	drirc: whitelist glthread for Euro Truck Simulator 2	2017-07-14 23:56:37 +02:00
Petr Sebor	483aa3997d	drirc: whitelist glthread for American Truck Simulator	2017-07-14 23:56:37 +02:00
Grigori Goronzy	d063168514	mesa/marshal: fix Windows build This was broken by commit `1ad24faa`. Reported by AppVeyor: https://ci.appveyor.com/project/mesa3d/mesa/build/4918	2017-07-14 23:46:21 +02:00
Edmondo Tommasina	952f21bc1e	drirc: whitelist glthread for The Witcher 2 Performance delta on AMD Phenom II X3 720 / RX 470 The Witcher 2: +18% Signed-off-by: Marek Olšák <marek.olsak@amd.com>	2017-07-14 23:41:59 +02:00
Edmondo Tommasina	9a9b9882a5	drirc: whitelist glthread for Civilization 5 Performance delta on AMD Phenom II X3 720 Civilization 5: +28% Signed-off-by: Marek Olšák <marek.olsak@amd.com>	2017-07-14 23:41:59 +02:00
Tim Rowley	818209118c	swr: JitManager runtime determination of architecture Fixes performance regression from `f50aa21456` - was forcing internal code generation to target AVX (no gather, etc). Reviewed-by: Bruce Cherniak <bruce.cherniak@intel.com>	2017-07-14 15:09:22 -05:00
Andres Gomez	25d43cd656	docs: update calendar, add news item and link release notes for 17.1.5 Signed-off-by: Andres Gomez <agomez@igalia.com>	2017-07-14 22:30:08 +03:00
Andres Gomez	7ad2f7078d	docs: add sha256 checksums for 17.1.5 Signed-off-by: Andres Gomez <agomez@igalia.com>	2017-07-14 22:30:08 +03:00
Andres Gomez	ea417c4c64	docs: add release notes for 17.1.5 Signed-off-by: Andres Gomez <agomez@igalia.com>	2017-07-14 22:30:08 +03:00
Grigori Goronzy	8d980bf920	st/mesa: Add KHR_no_error toggle to driconf Allows applications to be whitelisted. v2: Remove misguided DRI common part. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2017-07-14 21:23:44 +02:00
Grigori Goronzy	4909519a66	egl: Add EGL_KHR_create_context_no_error support This only adds the EGL side, needs to be plumbed into Mesa frontend. v2: Add check for extension availability. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2017-07-14 21:23:44 +02:00
Grigori Goronzy	2bbe235053	st/mesa: Add support for KHR_no_error flag Add a new context flag and plumb it through the various layers of the context creation code to set up dispatch tables for the no-error mode. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2017-07-14 21:23:40 +02:00
Grigori Goronzy	7299e82fa4	dri: Add KHR_no_error DRI extension This basic extension allows usage of the __DRI_CTX_FLAG_NO_ERROR flag. This includes support code for classic Mesa drivers to switch on the no-error mode if the flag is set. v2: Move to common DRI code. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2017-07-14 21:20:31 +02:00
Grigori Goronzy	cfbf60b0c2	mesa/marshal: fix glNamedBufferData with NULL data The semantics are similar to glBufferData. Tested-by: Marc Dietrich <marvin24@gmx.de> Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2017-07-14 21:20:31 +02:00
Grigori Goronzy	1ad24faa11	mesa/marshal: add marshalling for glClearBuffer* Add async marshalling/unmarshalling for all glClearBuffer variants. These entry points are commonly used in general and Alien Isolation specifically uses glClearBufferiv. Slightly reduces the number of thread synchronizations with glthread in that game. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2017-07-14 21:20:31 +02:00
Grigori Goronzy	8036198c0f	mesa/marshal: extract ClearBuffer helpers Extract clear buffer helper functions in preparation for adding marshal/unmarshal functions for the various glClearBuffer variants. v2: Fix command size. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2017-07-14 21:20:08 +02:00
Christoph Haag	98514e9959	gallium/hud: use double values for all graphs The fps graph for example calculates the fps as double with small variations based on when query_new_value() is called, which causes many values to be truncated on the cast to uint64_t. The HUD internally stores the values as double, so just use double everywhere instead of fixing this with rounding. Using doubles also allows the hud to show small variations instead of being clamped to discrete values. v2: Don't print decimals in the dump file when not necessary Signed-off-by: Christoph Haag <haagch+mesadev@frickel.club> Signed-off-by: Marek Olšák <marek.olsak@amd.com>	2017-07-14 17:34:39 +02:00
Lucas Stach	7e426ef6ec	Revert "etnaviv: add support for snorm textures" This reverts commit `d8b2ccdb88`, which causes priglit regressions on GPUs with SNORM support. We'll have another try at enabling this feature after the 17.2 branchpoint. Signed-off-by: Lucas Stach <l.stach@pengutronix.de>	2017-07-14 17:21:50 +02:00
Wladimir J. van der Laan	1d05cec205	etnaviv: reset indexed rendering information when not rendering indexed A dangling bo object would result in memory corruption while loading a level in ioquake3_opengl2. Fixes: `330d0607ed` (gallium: remove pipe_index_buffer and set_index_buffer) Suggested-by: Lucas Stach <l.stach@pengutronix.de> Signed-off-by: Wladimir J. van der Laan <laanwj@gmail.com> Reviewed-by: Lucas Stach <l.stach@pengutronix.de>	2017-07-14 17:19:42 +02:00
Wladimir J. van der Laan	bb2498a7f6	etnaviv: Use the correct LOG instruction on GC3000 GC3000 has a new LOG instruction, similar to the new SIN and COS instructions. Generate the new instruction sequence when appropriate; there are two occasions, as part of LIT and the generator for the LG2 instruction itself. Signed-off-by: Wladimir J. van der Laan <laanwj@gmail.com> Reviewed-by: Lucas Stach <l.stach@pengutronix.de>	2017-07-14 17:15:41 +02:00
Lucas Stach	bccd21ee88	etnaviv: flush source TS before resolve If we blit from a rendertarget or a depthstencil buffer there might still be dirty data in the TS buffer which needs to be flushed out. Fixes missing shadow tiles in glmark2 shadow. Signed-off-by: Lucas Stach <l.stach@pengutronix.de> Reviewed-by: Philipp Zabel <p.zabel@pengutronix.de>	2017-07-14 17:13:12 +02:00
Philipp Zabel	e9b3381715	etnaviv: flush color cache and depth cache together before resolves Before resolving a rendertarget or a depth/stencil resource into a texture, flush both the color cache and the depth cache together. It is unclear whether this is necessary for the following stall to work properly, or whether the depth flush just adds enough time for the color cache flush to finish before the resolver is started, but this change removes artifacts that otherwise appear if a texture is sampled directly after rendering into it. The test case is a simple QML scene graph with a QtWebEngine based WebView rendered on top of a blue background: import QtQuick 2.0 import QtQuick.Window 2.2 import QtWebView 1.1 Window { Rectangle { id: background anchors.fill: parent color: "blue" } WebView { id: webView anchors.fill: parent } Component.onCompleted: { webView.url = "<some animated website>" } } If the website is animated, the WebView renders the site contents into texture tiles and immediately afterwards samples from them to draw the tiles into the Qt renderbuffer. Without this patch, a small irregular triangle in the lower right of each browser tile appears solid blue, as if the texture sampler samples zeroes instead of the website contents, and the previously rendered blue Rectangle shows through. Other attempts such as adding a pipeline stall before the color flush or a TS cache flush afterwards or flushing multiple times, with stalls before and after each flush, have shown no effect. Signed-off-by: Philipp Zabel <p.zabel@pengutronix.de>	2017-07-14 17:12:36 +02:00
Lucas Stach	a98c1fbd9b	st/mesa: handle stfbi being NULL on entry of st_framebuffer_reuse_or_create Apparently this can happen. Just bail out early in that case, as all the called functions return NULL in that case. Fixes weston-terminal for me. Fixes: `147d7fb772` ("st/mesa: add a winsys buffers list in st_context") Signed-off-by: Lucas Stach <l.stach@pengutronix.de> Reviewed-by: Charmaine Lee <charmainel@vmware.com>	2017-07-14 17:12:00 +02:00
Daniel Stone	5295df63ad	egl/wayland: Use MIN2 for wl_drm version Use a slightly more explicit version cap for binding wl_drm, so we can add other interfaces with different versioning schemes later. Reviewed-by: Eric Engestrom <eric.engestrom@imgtec.com> Reviewed-by: Emil Velikov <emil.velikov@collabora.com>	2017-07-14 14:14:05 +01:00
Daniel Stone	4b8ef27e84	egl/wayland: Fix whitespace damage Convert tabs to spaces, fix misalignments. Reviewed-by: Eric Engestrom <eric.engestrom@imgtec.com> Reviewed-by: Emil Velikov <emil.velikov@collabora.com>	2017-07-14 14:14:05 +01:00
Daniel Stone	2b895475f6	util: Remove u_math from u_vector u_vector.h doesn't actually use anything from u_math, but it does mean everyone has to pull in src/gallium/auxiliary/util includes. Just remove it, adding a <string.h> include to u_vector.c to cover memcpy. Reviewed-by: Eric Engestrom <eric.engestrom@imgtec.com> Reviewed-by: Emil Velikov <emil.velikov@collabora.com>	2017-07-14 14:14:05 +01:00
Eric Engestrom	8821ef4be1	configure: only install khrplatform.h if needed khrplatform.h is only used by EGL and GLES; let's only install it when one of those is enabled. Cc: mesa-stable@lists.freedesktop.org Signed-off-by: Eric Engestrom <eric.engestrom@imgtec.com> Reviewed-by: Jussi Kukkonen <jussi.kukkonen@intel.com> Reviewed-by: Emil Velikov <emil.velikov@collabora.com>	2017-07-14 13:23:54 +01:00
Eric Engestrom	b50b4b6f84	scons: split out check_header() helper Signed-off-by: Eric Engestrom <eric.engestrom@imgtec.com>	2017-07-14 13:23:54 +01:00
Juan A. Suarez Romero	5cd4ece34e	anv/pipeline: do not use BITFIELD64_BIT() In the previous commit, forgot to apply v2 suggestions. Fixes: `28d0c38` (anv/pipeline: use unsigned long long constant to check enable vertex inputs) Signed-off-by: Juan A. Suarez Romero <jasuarez@igalia.com>	2017-07-14 10:33:19 +00:00
Juan A. Suarez Romero	28d0c38d85	anv/pipeline: use unsigned long long constant to check enable vertex inputs When initializing the ANV pipeline, one of the tasks is checking which vertex inputs are enabled. This is done by checking if the enabled bits in inputs_read. But the mask to use is computed doing `(1 << (VERT_ATTRIB_GENERIC0 + desc->location))`. The problem here is that if location is 15 or greater, the sum is 32 or greater. But C is handling 1 as a 32-bit integer, which means the displaced bit is out of range and thus the full value is 0. Thus, use 1ull, which is an unsigned long long value. This fixes: dEQP-VK.pipeline.vertex_input.max_attributes.16_attributes.binding_one_to_one.interleaved v2: use 1ull instead of BITFIELD64_BIT() (Matt Turner) Reviewed-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> Reviewed-by: Matt Turner <mattst88@gmail.com> Signed-off-by: Juan A. Suarez Romero <jasuarez@igalia.com> Cc: mesa-stable@lists.freedesktop.org	2017-07-14 08:09:18 +00:00
Kenneth Graunke	b2da123801	i965: Use pushed UBO data in the scalar backend. This actually takes advantage of the newly pushed UBO data, avoiding pull loads. Improves performance in GLBenchmark Manhattan 3.1 by: HSW: ~1%, BDW/SKL/KBL GT2: 3-4%, SKL GT4: 7-8%, APL: 4-5%. (thanks to Eero Tamminen for these numbers) shader-db results on Skylake, ignoring programs with spill/fill changes: total instructions in shared programs: 13963994 -> 13651893 (-2.24%) instructions in affected programs: 4250328 -> 3938227 (-7.34%) helped: 28527 HURT: 0 total cycles in shared programs: 179808608 -> 172535170 (-4.05%) cycles in affected programs: 79720410 -> 72446972 (-9.12%) helped: 26951 HURT: 1248 LOST: 46 GAINED: 21 Many "Deus Ex: Mankind Divided" shaders which already spilled end up spill a lot more (about 240 programs hurt, 9 helped). The cycle estimator suggests this is still overall a win (-0.23% in cycle counts) presumably because we trade pull loads for fills. v2: Drop "PULL" environment variable left in for initial debugging (caught by Matt). Reviewed-by: Matt Turner <mattst88@gmail.com>	2017-07-13 20:18:54 -07:00
Kenneth Graunke	c9ef27e77b	i965: Factor out push locations. With UBOs, the answer of "have we decided to push this uniform" gets a bit more complicated - for one, we have multiple surfaces. This patch refactors things so we can add the new code in a single place. Reviewed-by: Matt Turner <mattst88@gmail.com>	2017-07-13 20:18:54 -07:00

1 2 3 4 5 ...

94003 Commits All Branches Search

94003 Commits

All Branches